Score
Designs and writes production-quality code for neural network models and components, including architectures, layer and kernel implementations, and training/inference hooks; integrates these into deep learning frameworks and maintains compatibility across framework versions. Builds performant, maintainable implementations (e.g., custom layers, regularization kernels, spectral-aware convolutions, optimization hooks) while minimizing inference overhead.
The choice between PyTorch and TensorFlow remains a critical decision for AI researchers and practitioners, yet systematic, empirically grounded comparisons across usability, training/inference performance, and production deployment capabilities are lacking. Method: We conduct a comprehensive benchmarking study—including XLA, TensorRT, and other backend accelerators—analyze code complexity, evaluate cross-framework interoperability (ONNX, TorchScript, TFLite), and survey state-of-the-art literature and ecosystem tooling. Contribution/Results: Our analysis reveals fundamental paradigmatic differences: PyTorch’s dynamic computation graph excels in research agility and prototyping flexibility, whereas TensorFlow’s static graph design delivers superior end-to-end deployment maturity, multi-platform support (e.g., mobile, edge), and enterprise service integration. Computationally, both frameworks achieve comparable peak performance; however, their ecosystem roles have significantly diverged. We identify cross-framework interoperability and unified compiler-level optimization as pivotal future directions, providing evidence-based guidance for framework selection in AI development.
Neural network migration across mainstream frameworks (e.g., PyTorch and TensorFlow) remains challenging due to manual reconstruction requirements, poor compatibility, and semantic discrepancies. To address this, we propose a fully automated cross-framework migration method based on a hub-style intermediate representation (Hub IR). Our approach constructs a unified model IR via abstract syntax tree parsing, then performs semantic-aware structural mapping and framework-specific code generation to achieve bidirectional, functionally equivalent model translation. We systematically resolve two core challenges: cross-framework semantic divergence and topological structure mismatch—addressed for the first time in a unified framework. Experimental evaluation on five representative neural networks demonstrates functional equivalence of generated code, over 90% reduction in manual intervention, and substantial improvements in migration reliability and development efficiency.
Traditional software engineering design principles—particularly SOLID—are often assumed to apply uniformly across domains, yet their applicability and interpretation in AI framework design remain underexplored. Method: This study conducts a systematic, context-sensitive evaluation of TensorFlow and scikit-learn against SOLID principles through architectural documentation analysis, source-code inspection, and comparative design philosophy assessment, yielding a five-dimensional principle-mapping framework. Contribution/Results: We demonstrate that neither framework strictly adheres to nor violates SOLID; rather, both dynamically prioritize principles based on AI-specific constraints—e.g., experimental iteration, computational efficiency, and maintainability. TensorFlow emphasizes performance at the expense of Single Responsibility and Interface Segregation, while scikit-learn aligns more closely with SOLID overall but makes localized efficiency-driven compromises in critical paths. Crucially, we introduce the “domain-aware design principle evolution paradigm,” arguing that AI frameworks require an interpretable architectural trade-off model—one that explicitly reconciles rigorous software engineering principles with pragmatic AI development needs.
Contemporary machine learning research overemphasizes model performance while neglecting resource efficiency and long-term sustainability. Method: This paper systematically identifies memory/GPU-memory leak–inducing code smells in PyTorch, TensorFlow, and Keras—derived from developer community discussions and real-world code snippets—yielding 30 PyTorch-specific and 16 TensorFlow/Keras-specific anti-patterns. It proposes the first cross-framework resource leakage taxonomy, balancing generality with framework-specific adaptability, validated through a three-stage empirical process: qualitative analysis, normative refinement, and practical evaluation. Contribution/Results: The study yields 50 actionable best practices for mitigating resource leaks. It fills a critical gap in ML engineering research on resource efficiency and provides a structured, methodology-grounded coding guide for building efficient, robust, and sustainable machine learning systems.
Deep learning (DL) code refactoring lacks systematic investigation, and existing IDEs and refactoring tools lack support for DL-specific semantics—such as tensor operations and automatic differentiation. Method: We conduct the first large-scale empirical study, analyzing 4,921 refactoring commits across five mainstream DL projects (e.g., PyTorch) and surveying 159 practitioners. Using manual commit analysis, experience mining, and cross-project statistical comparison, we characterize DL refactoring patterns and tooling gaps. Contribution/Results: We find that DL refactoring predominantly targets model architecture and data pipeline adjustments—differing significantly from traditional Java software in type distribution. Current tools universally lack DL semantic awareness. Based on these findings, we propose design principles for DL-aware refactoring tools, emphasizing tensor dependency modeling and computational graph awareness. We further formulate a practical, actionable roadmap for integrating these capabilities into next-generation DL development environments.
This paper addresses the challenge of accurately identifying parallelization opportunities in complex loops using static analysis. To tackle this, we propose a deep learning–based code parallelism prediction framework. Methodologically, we design a genetic algorithm to automatically generate diverse loop code samples—covering both clearly parallelizable cases and those with ambiguous data dependencies—and construct a manually annotated training dataset. We then employ both deep neural networks (DNNs) and convolutional neural networks (CNNs) to model and classify tokenized code sequences. Experimental results show that CNNs achieve marginally higher average accuracy, while both models demonstrate robust performance. Our key contributions are threefold: (1) the first integration of generative genetic algorithms with deep learning for parallelism prediction; (2) effective mitigation of data scarcity and ambiguity in dependency analysis; and (3) empirical validation that training data diversity critically enhances model generalization—establishing a novel paradigm for automated parallel optimization.
This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.
This work addresses the significant performance degradation of existing large code models in industrial settings characterized by strong hardware semantics, domain-specific language structures, and stringent resource constraints. To bridge this gap, we propose the first industrial-scale unified code foundation model with 32 billion parameters, spanning critical domains including chip design, GPU kernel optimization, embedded systems, compiler optimization, and 3D modeling. The model is trained from scratch, integrating general-purpose code pretraining, curated industrial code annealing, progressive long-context expansion from 8K to 128K tokens, and execution-based post-training strategies. Experimental results demonstrate competitive performance across 14 general-purpose benchmarks and establish state-of-the-art open-source baselines on nine industrial benchmarks across four key domains.
This work addresses the limitations of existing automated documentation generation methods, which often overlook structural and quantitative code features that critically influence code readability, resulting in contextually irrelevant or inaccurate documentation in computational notebooks. To bridge this gap, the study systematically introduces code metrics as auxiliary signals for the first time, constructing a high-quality dataset of (code, Markdown) pairs and integrating metric information into both CNN-RNN and GPT-3.5 architectures. Experimental results demonstrate significant improvements in generation quality: the CNN-RNN model achieves a 6% increase in BLEU-1 and a 3% gain in ROUGE-L F1, while few-shot GPT-3.5 shows a 9% improvement in BERTScore F1. These findings validate the generalizability and effectiveness of code metrics as enhancing signals across diverse model paradigms.
This work addresses the inefficiency and high cost of migrating deep learning models across frameworks—such as from TensorFlow to JAX—in large-scale AI systems. To tackle this challenge, the authors propose an automated, multi-agent collaborative migration approach that integrates static code analysis with an AI-driven planner to generate precise migration instructions. A coordinator and encoder work in tandem, leveraging AI-generated, example-driven migration guides to achieve high-fidelity translation without requiring test code. Innovatively, an AI-based evaluator assesses migration quality, establishing a self-reinforcing development loop. Evaluated in real-world, large-scale production environments, the method accelerates framework migration by 6.4–8×, substantially expediting model infrastructure evolution.