design efficient backbones

Design and evaluate neural network backbone architectures, specifying layer and block types and arrangements to balance parameter count and accuracy, maximize throughput on target hardware, and minimize memory and compute footprint.

designefficientbackbones

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Nov 20, 2024
JF
Jared Fernandez
🏛️ Meta | Carnegie Mellon University

Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.

Assessing diminishing returns in scaling accelerators for large model trainingEvaluating parallelization strategies to minimize distributed communication overheadOptimizing hardware configuration for efficient large-scale model training

This work proposes a hardware-aware, extensible neural architecture search (NAS) framework that decouples the search, optimization, and deployment pipelines to overcome the tight coupling prevalent in existing NAS methods, which hinders adaptation to new hardware or custom operators. The framework employs YAML for unified search space specification and leverages Optuna for efficient architecture search. It supports hardware-in-the-loop evaluation by integrating Docker-based cross-compilation, automated on-device binary generation, and multidimensional performance metrics—including FLOPs, parameter count, and latency—thereby significantly enhancing support for heterogeneous accelerator platforms and reducing engineering overhead in embedded AI deployment.

deploymentextensibilityhardware-aware

Neural architecture design often relies on heuristic rules or expensive search strategies, lacking a principled, differentiable mapping from performance to structure. Method: This paper proposes an automatic architecture optimization framework grounded in structure–performance mapping modeling. We introduce the Architecture Synthesis Neural Network (ASNN), the first model that takes a performance distribution (e.g., accuracy distribution) as input and differentiably synthesizes high-performing architectural parameters—enabling invertible, generalizable mapping from performance to structure. Leveraging a TensorFlow-based multi-layer network performance dataset, ASNN learns this inverse mapping via neural regression and incorporates an iterative prediction mechanism for progressive refinement. Contribution/Results: On both two- and three-layer networks, ASNN discovers novel architectures surpassing the original dataset’s best-performing models, achieving statistically significant average test accuracy improvements. Experiments validate its effectiveness in architecture recommendation, cross-architecture generalization, and iterative optimization.

ASNN automates neural network design efficientlyASNN learns relationship between NN architecture and accuracyASNN suggests improved architectures for better performance

Survey on Characterizing and Understanding GNNs from a Computer Architecture Perspective

Aug 04, 2024
MW
Meng Wu
🏛️ Chinese Academy of Sciences | University of Chinese Academy of Sciences | Hong Kong University of Science and Technology

Graph Neural Networks (GNNs) suffer from low execution efficiency and poor training stability on parallel and distributed systems due to unoptimized computational, memory access, and communication patterns. Method: This paper adopts a computer architecture–centric perspective to systematically analyze these patterns and identify critical performance bottlenecks. It proposes the first hardware-software co-design–oriented, three-level GNN workload taxonomy—unifying architectural characterization across diverse GNN models. Leveraging architecture-aware analysis, fine-grained performance modeling, and cross-platform empirical evaluation, it quantifies behavioral disparities of GNN workloads across CPUs, GPUs, and distributed clusters. Contribution/Results: The work delivers a reusable modeling framework and actionable optimization guidelines for efficient, architecture-agnostic GNN implementation. It significantly improves hardware resource utilization and training stability while enabling systematic, principled design of high-performance GNN systems.

Graph Neural NetworksParallel ProcessingPerformance Optimization

Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision

Jun 09, 2024
PJ
Pranav Jeevan
🏛️ Indian Institute of Technology Bombay

Lightweight backbone networks lack systematic evaluation in few-shot cross-domain vision tasks. Method: Under a unified training protocol, this work conducts the first cross-domain, cross-data-scale comparative study—spanning natural (ImageNet), medical (CheXpert), astronomical (SDSS), and remote sensing (EuroSAT) domains—evaluating ConvNeXt, RegNet, EfficientNet, and ViT architectures under low-data fine-tuning regimes. Results: CNN-based backbones—particularly ConvNeXt and RegNet—consistently outperform ViT variants under limited data, demonstrating superior generalization and cross-domain robustness. Ablation analysis uncovers principled links between architectural design choices (e.g., locality bias, inductive priors) and domain-specific data characteristics (e.g., spatial coherence, label scarcity). The study releases an open-source benchmarking framework and standardized results, providing empirically grounded guidance for backbone selection in resource-constrained settings.

Compares performance of backbones on small datasets.Evaluates lightweight CNN backbones across diverse domains.Identifies best backbones for specific computer vision tasks.

Latest Papers

What's happening recently
View more

Existing hardware-software co-design tools struggle to accurately model memory consumption and backward-pass complexity in neural network training. This work proposes the first extension of the experimentally validated inference modeling framework, Stream, to the training domain, introducing a comprehensive framework for modeling and optimizing training on heterogeneous dataflow accelerators. The framework supports training workflow modeling, exploration of layer fusion configurations, and optimization of activation checkpointing strategies. Integrated with a genetic algorithm for hardware architecture search, it is validated on ResNet-18 and a small-scale GPT-2 model, effectively uncovering critical trade-offs between performance and memory in training-specific hardware design and identifying superior architectures and training strategies.

backpropagation complexityhardware-software co-designheterogeneous accelerators

This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.

deep learning librarieseducational gapfundamental understanding

This work addresses the significant discrepancy between traditional MACs-based efficiency metrics for vision backbones and actual inference latency on edge devices, which hinders hardware-efficient design. By analyzing the divergence between theoretical MACs and real-world execution times of common building blocks, the study identifies key factors governing hardware efficiency. It proposes LowFormer, a novel backbone featuring the lightweight Lowtention module as a replacement for multi-head self-attention. Through hardware-aware co-design of macro- and micro-architectures alongside cross-platform deployment optimizations, LowFormer achieves higher ImageNet accuracy while substantially outperforming state-of-the-art models in speed across diverse hardware platforms—including both edge and desktop GPUs—and demonstrates strong performance on downstream tasks such as classification, detection, and segmentation.

edge devicesexecution timehardware efficiency

Hot Scholars

JT

Jin Tang

Anhui University
Computer visionintelligent video analysis
CC

Cheng Cui

BUAA
deep learningnetwork designOCRmllm
ST

Siliang Tang

Professor of Computer Science, Zhejiang University
Natural Language ProcessingCross-media AnalysisGraph Neural Network
LS

Linlin Shen

Shenzhen University
Deep LearningComputer VisionFacial Analysis/RecognitionMedical Image Analysis
JW

Jun-Wei Hsieh

National Yang Ming Chiao Tung University
computer visionAIimage processing