Score
Designs, implements, configures, or evaluates software libraries and runtime systems that support development, training, debugging, and deployment of deep neural network models, including core components such as computation graphs, automatic differentiation, optimizers, data pipelines, model serialization, hardware acceleration, and distributed training. Also includes building APIs, tooling for model inspection and profiling, and extending frameworks to add new layers, operators, or execution backends.
The choice between PyTorch and TensorFlow remains a critical decision for AI researchers and practitioners, yet systematic, empirically grounded comparisons across usability, training/inference performance, and production deployment capabilities are lacking. Method: We conduct a comprehensive benchmarking study—including XLA, TensorRT, and other backend accelerators—analyze code complexity, evaluate cross-framework interoperability (ONNX, TorchScript, TFLite), and survey state-of-the-art literature and ecosystem tooling. Contribution/Results: Our analysis reveals fundamental paradigmatic differences: PyTorch’s dynamic computation graph excels in research agility and prototyping flexibility, whereas TensorFlow’s static graph design delivers superior end-to-end deployment maturity, multi-platform support (e.g., mobile, edge), and enterprise service integration. Computationally, both frameworks achieve comparable peak performance; however, their ecosystem roles have significantly diverged. We identify cross-framework interoperability and unified compiler-level optimization as pivotal future directions, providing evidence-based guidance for framework selection in AI development.
Neural networks frequently exhibit hard-to-diagnose and hard-to-fix anomalous behaviors in production, yet existing maintenance tools are heavily skewed toward the training phase and provide inadequate support for post-deployment diagnosis and root-cause analysis. Method: We adopt a qualitative research approach, conducting in-depth interviews and complementary surveys with 23 practitioners to systematically characterize real-world challenges and tooling gaps in neural network testing, debugging, and maintenance. Contribution/Results: Our study is the first to empirically identify critical shortcomings in current tooling—particularly the neglect of runtime anomaly interpretation, error attribution, and repair validation. We find practitioners urgently require a new maintenance paradigm centered on behavioral observability, causal reasoning, and iterative repair. These findings provide empirical grounding and actionable design principles for building next-generation neural network maintenance infrastructures.
Deep learning (DL) systems exhibit strong coupling between data and models, impeding conventional unit testing; existing software engineering research has not systematically addressed this challenge. This paper proposes *Simulated Deep Testing*, a novel methodology that enables unit-level independent testing of DL data preparation and model design through workflow decoupling, modular architecture, and dependency mocking. Our key contributions are: (1) the first mock-based unit testing paradigm specifically designed for DL applications; (2) principled guidelines for decoupling DL workflows; and (3) KUnit—the first mock-testing framework supporting Keras. Evaluated on 50 real-world DL projects, KUnit detected 63 defects; a user study confirmed that developers successfully resolved all identified issues using KUnit, demonstrating its effectiveness in early defect detection and lightweight component validation.
Deep learning (DL) code refactoring lacks systematic investigation, and existing IDEs and refactoring tools lack support for DL-specific semantics—such as tensor operations and automatic differentiation. Method: We conduct the first large-scale empirical study, analyzing 4,921 refactoring commits across five mainstream DL projects (e.g., PyTorch) and surveying 159 practitioners. Using manual commit analysis, experience mining, and cross-project statistical comparison, we characterize DL refactoring patterns and tooling gaps. Contribution/Results: We find that DL refactoring predominantly targets model architecture and data pipeline adjustments—differing significantly from traditional Java software in type distribution. Current tools universally lack DL semantic awareness. Based on these findings, we propose design principles for DL-aware refactoring tools, emphasizing tensor dependency modeling and computational graph awareness. We further formulate a practical, actionable roadmap for integrating these capabilities into next-generation DL development environments.
Deep learning (DL) infrastructure libraries—including frameworks, compilers, and hardware-accelerator libraries—lack systematic definitions and methodological foundations for testing. Method: This paper presents the first comprehensive study on DL library testing: (1) it proposes a novel three-category taxonomy of DL libraries and corresponding defect definitions, unifying the scope and boundaries of DL library defects and testing; (2) it introduces the first end-to-end testing methodology framework covering the full DL stack, with a structured survey of domain-specific testing techniques and tools; and (3) it systematically identifies six cross-cutting testing challenges and outlines concrete directions for future research. Contribution/Results: Based on rigorous literature review, taxonomic analysis, and comparative evaluation, this work establishes the first authoritative technical guideline for DL library testing—providing industry practitioners with actionable, deployable testing strategies and offering academia a benchmark paradigm for further investigation.
Increasing complexity in AI systems has led to high code redundancy, poor reusability, and escalating maintenance costs. Method: This paper proposes the first systematic object-oriented programming (OOP) mapping framework tailored for AI engineering practice. It deeply integrates core OOP principles—encapsulation, inheritance, and polymorphism—across the entire ML/DL/LLM pipeline (data preprocessing, model training, evaluation, and deployment), augmented by design patterns such as Factory and Strategy to construct a reusable AI component library and modular architecture. Contribution/Results: The framework introduces a semantic modeling methodology that formally aligns OOP principles with AI workflows and provides native Python support. Empirical evaluation demonstrates over 30% reduction in code redundancy and significantly improved cross-project component reuse, thereby enabling maintainable, iterative development of industrial-scale AI systems.
This work addresses the lack of publicly available, diverse benchmark datasets for systematically evaluating neural network code verification, refactoring, and migration tools. To bridge this gap, the authors propose a novel approach that leverages large language models to automatically generate neural network code spanning a wide range of architectural components, input types, and tasks. The generated samples are rigorously validated through static analysis and symbolic tracing to ensure both structural and semantic adherence to precise design specifications. The resulting benchmark comprises 608 correct and diverse neural network implementations, constituting the first publicly reusable dataset of its kind. This resource significantly advances reproducibility and enables systematic evaluation in research on neural network reliability and maintainability.
This work addresses the challenge of defect detection in deep learning libraries such as TensorFlow and PyTorch, where complex APIs often lead to subtle bugs and existing testing approaches suffer from high false-positive rates due to imprecise specifications. To overcome this limitation, the authors propose a machine learning classifier that leverages tensor shape abstraction as a precise input representation for API validity constraints. By integrating runtime feedback to automatically generate labeled training data, the method learns accurate usage patterns without relying on manual annotations. Implemented within the ACETest framework, the approach achieves over 91% classification accuracy across 183 APIs and significantly improves test pass rates—from 29% to 61%—demonstrating enhanced precision and scalability in testing deep learning libraries.
This work addresses the computational, memory, and storage bottlenecks associated with deploying large-scale deep neural networks (DNNs) in resource-constrained environments. The authors propose a novel pruning method that integrates system-level engineering requirements with human-interpretable concepts—such as color and semantic categories—to identify critical neurons through analysis of their activation patterns, thereby guiding the generation of lightweight models. Notably, this approach is the first to incorporate interpretable concepts directly into the DNN pruning pipeline. Evaluated on VGG-19 using a dataset comprising 26,384 RGB images, the method yields pruned models that achieve substantial reductions in model size and computational overhead while maintaining high performance, demonstrating strong applicability across diverse real-world scenarios with stringent resource constraints.
Existing hardware-software co-design tools struggle to accurately model memory consumption and backward-pass complexity in neural network training. This work proposes the first extension of the experimentally validated inference modeling framework, Stream, to the training domain, introducing a comprehensive framework for modeling and optimizing training on heterogeneous dataflow accelerators. The framework supports training workflow modeling, exploration of layer fusion configurations, and optimization of activation checkpointing strategies. Integrated with a genetic algorithm for hardware architecture search, it is validated on ResNet-18 and a small-scale GPT-2 model, effectively uncovering critical trade-offs between performance and memory in training-specific hardware design and identifying superior architectures and training strategies.
This study addresses a critical yet underexplored class of defects—referred to as fBugs—in the frontends of deep learning compilers during the conversion of programs into graph-based intermediate representations. Focusing on TorchDynamo, the default frontend of PyTorch 2, this work presents the first systematic empirical investigation of fBugs by leveraging a domain-knowledge-enhanced large language model to analyze 123 real-world bugs. The authors establish a comprehensive taxonomy encompassing seven root-cause categories and fifteen subcategories, and develop root-cause-aware test cases. Moving beyond conventional black-box or low-level API–centric approaches, their methodology successfully uncovers 23 previously unknown fBugs—15 of which have been confirmed—spanning eight subcategories, thereby substantially improving the robustness of compiler frontends.