Score
Designs, implements, and evaluates deep neural network architectures and models—including convolutional networks and large-scale networks—along with the software frameworks and training pipelines that support them. Builds and tunes optimization and training procedures, conducts research and experiments on methods and architectures, and integrates trained models into larger systems while addressing performance, scalability, and reliability.
Existing deep learning surveys suffer from incomplete coverage, overemphasis on CNNs, and a lack of cross-domain comparative analysis of model efficacy. Method: We systematically survey state-of-the-art (SOTA) models across five domains—computer vision, natural language processing, time-series analysis, ubiquitous computing, and robotics—encompassing Transformer, CNN, RNN, GNN, self-supervised, and multimodal learning architectures. Contribution/Results: We propose the first unified survey framework spanning multiple domains while integrating theoretical foundations with practical problem-solving capabilities. We establish a cross-domain model capability comparison system that clarifies principled criteria for optimal model selection per scenario. Furthermore, we identify scalability, robustness, and energy efficiency as three fundamental, domain-agnostic challenges—offering concrete directions for future research. This work bridges critical gaps in both scope and analytical depth, enabling more informed architectural choices and fostering cross-disciplinary innovation.
High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.
Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.
Traditional software engineering design principles—particularly SOLID—are often assumed to apply uniformly across domains, yet their applicability and interpretation in AI framework design remain underexplored. Method: This study conducts a systematic, context-sensitive evaluation of TensorFlow and scikit-learn against SOLID principles through architectural documentation analysis, source-code inspection, and comparative design philosophy assessment, yielding a five-dimensional principle-mapping framework. Contribution/Results: We demonstrate that neither framework strictly adheres to nor violates SOLID; rather, both dynamically prioritize principles based on AI-specific constraints—e.g., experimental iteration, computational efficiency, and maintainability. TensorFlow emphasizes performance at the expense of Single Responsibility and Interface Segregation, while scikit-learn aligns more closely with SOLID overall but makes localized efficiency-driven compromises in critical paths. Crucially, we introduce the “domain-aware design principle evolution paradigm,” arguing that AI frameworks require an interpretable architectural trade-off model—one that explicitly reconciles rigorous software engineering principles with pragmatic AI development needs.
This study addresses the lack of systematic understanding regarding the trade-off between computational efficiency and model accuracy in convolutional neural networks (CNNs) under distributed training settings. It presents the first comprehensive analysis of how different CNN architectures and data augmentation strategies jointly influence model accuracy and resource consumption in such scenarios. Through extensive comparative experiments, the work evaluates the performance and computational overhead of various architecture–augmentation combinations. The findings reveal that specific pairings can significantly enhance training efficiency while preserving model accuracy, offering both theoretical insights and practical guidance for deploying efficient models in resource-constrained environments.
Designing deep learning accelerators for heterogeneous HPC and edge platforms faces key challenges including insufficient parallelism exploitation and excessive data movement overhead. This paper systematically surveys accelerator design methodologies, covering hardware-software co-design, high-level synthesis, domain-specific compilers (e.g., TVM, Halide), design space exploration, and cycle-accurate modeling and simulation. We propose, for the first time, a unified multi-dimensional classification framework that distills two fundamental principles: “minimizing data movement” and “maximizing parallelism.” The survey bridges the gap between architectural overviews and implementation-oriented methodologies, explicitly identifying emerging directions such as approximate computing integrated with reconfigurability. Our work provides both a methodological foundation and practical guidance for developing efficient, scalable AI accelerators—enabling principled design decisions across diverse heterogeneous computing ecosystems.
Existing hardware-software co-design tools struggle to accurately model memory consumption and backward-pass complexity in neural network training. This work proposes the first extension of the experimentally validated inference modeling framework, Stream, to the training domain, introducing a comprehensive framework for modeling and optimizing training on heterogeneous dataflow accelerators. The framework supports training workflow modeling, exploration of layer fusion configurations, and optimization of activation checkpointing strategies. Integrated with a genetic algorithm for hardware architecture search, it is validated on ResNet-18 and a small-scale GPT-2 model, effectively uncovering critical trade-offs between performance and memory in training-specific hardware design and identifying superior architectures and training strategies.
This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.
Distributed deep learning (DDL) frameworks suffer from a lack of systematic understanding of defects, hindering robustness and maintainability. Method: We conduct the first large-scale empirical study of 849 real-world issues across DeepSpeed, Megatron-LM, and Colossal-AI, proposing the first defect taxonomy tailored to specialized DDL frameworks—comprising 34 symptom categories, 28 root-cause categories, and 6 repair patterns—and establishing a phase-aware symptom–cause–repair mapping across six execution stages: initialization, communication, computation, memory, fault tolerance, and scheduling. Results: We find that 45.1% of symptoms—and 95% of communication-related configuration issues—are uniquely distributed in nature; over 60% of defects are resolvable via version/dependency management or distributed tuning. Setup failures, memory anomalies, and performance deviations emerge as the top three distributed-specific defect types. We quantify root-cause distributions per stage and distill reusable repair patterns and engineering best practices, directly supporting enhanced framework robustness.
This work addresses the lack of publicly available, diverse benchmark datasets for systematically evaluating neural network code verification, refactoring, and migration tools. To bridge this gap, the authors propose a novel approach that leverages large language models to automatically generate neural network code spanning a wide range of architectural components, input types, and tasks. The generated samples are rigorously validated through static analysis and symbolic tracing to ensure both structural and semantic adherence to precise design specifications. The resulting benchmark comprises 608 correct and diverse neural network implementations, constituting the first publicly reusable dataset of its kind. This resource significantly advances reproducibility and enables systematic evaluation in research on neural network reliability and maintainability.
Traditional small-scale datasets such as MNIST struggle to effectively differentiate the performance of advanced neural network architectures, particularly due to their limited capacity to capture inductive biases for sequential data. This work presents the first systematic evaluation of diverse models—including ResNet, Temporal Convolutional Networks (TCNs), and Dilated Convolutional Neural Networks (DCNNs)—on the lightweight, structured sequential dataset MNIST-1D, benchmarking them against baselines such as logistic regression, MLPs, CNNs, and GRUs. Experimental results demonstrate that TCNs and DCNNs substantially outperform conventional approaches, achieving near-human accuracy, while ResNet also exhibits strong performance. These findings validate MNIST-1D as an efficient and effective benchmark and underscore the critical role of inductive bias in resource-constrained settings.