custom layer development

Designs and implements custom neural-network layer modules for deep learning frameworks, including their forward and backward computations, parameter and state management, shape/dtype checks, integration with automatic-differentiation APIs, serialization, and unit tests; may also implement optimized CPU/GPU kernels and performance/compatibility validation for the new layer.

customlayerdevelopment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.

deep learning librarieseducational gapfundamental understanding

Neural Network Interoperability Across Platforms

Nov 04, 2025
ND
Nadia Daoudi
🏛️ Luxembourg Institute of Science and Technology | University of Luxembourg

Neural network migration across mainstream frameworks (e.g., PyTorch and TensorFlow) remains challenging due to manual reconstruction requirements, poor compatibility, and semantic discrepancies. To address this, we propose a fully automated cross-framework migration method based on a hub-style intermediate representation (Hub IR). Our approach constructs a unified model IR via abstract syntax tree parsing, then performs semantic-aware structural mapping and framework-specific code generation to achieve bidirectional, functionally equivalent model translation. We systematically resolve two core challenges: cross-framework semantic divergence and topological structure mismatch—addressed for the first time in a unified framework. Experimental evaluation on five representative neural networks demonstrates functional equivalence of generated code, over 90% reduction in manual intervention, and substantial improvements in migration reliability and development efficiency.

Addressing interoperability challenges across platforms like PyTorch and TensorFlowAutomating neural network migration between deep learning frameworksEliminating manual effort to modernize outdated neural network implementations

This work addresses the challenge of defect detection in deep learning libraries such as TensorFlow and PyTorch, where complex APIs often lead to subtle bugs and existing testing approaches suffer from high false-positive rates due to imprecise specifications. To overcome this limitation, the authors propose a machine learning classifier that leverages tensor shape abstraction as a precise input representation for API validity constraints. By integrating runtime feedback to automatically generate labeled training data, the method learns accurate usage patterns without relying on manual annotations. Implemented within the ACETest framework, the approach achieves over 91% classification accuracy across 183 APIs and significantly improves test pass rates—from 29% to 61%—demonstrating enhanced precision and scalability in testing deep learning libraries.

API specificationsbug detectionDeep Learning libraries

Deploying deep neural networks (DNNs) on resource-constrained hardware such as FPGAs poses a fundamental challenge in simultaneously optimizing performance, accuracy, and hardware resource utilization (e.g., DSPs and LUTs), while heavily relying on manual expertise. Method: This paper introduces the first unified, fully automated framework that integrates high-level synthesis (HLS) metaprogramming with programmable DNN optimization. It jointly models compiler optimizations, hardware mapping, and model compression, enabling customizable transformations and Bayesian-driven design-space exploration for end-to-end, cross-stage co-optimization. Contribution/Results: Experimental evaluation demonstrates that, without sacrificing original model accuracy, the framework reduces DSP and LUT usage by up to 92% and 89%, respectively, and accelerates optimization efficiency by 15.6× over exhaustive grid search—significantly diminishing reliance on domain expertise and manual tuning.

Automates DNN optimization for resource-constrained hardwareEnhances performance, accuracy, and resource efficiency on FPGAsReduces manual effort and optimization time significantly

NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models

Aug 15, 2025
XB
Xiaohan Bi
🏛️ Beihang University | Peng Cheng Laboratory | National University of Singapore | Hangzhou Innovation Institute of Beihang University

To address the challenges in DNN model reuse—namely, excessive inference overhead from full-model reuse and the limitations of existing Modular-with-Training (MwT) approaches, which support only small-scale CNNs and operate solely at the convolutional kernel granularity—this paper proposes NeMo, the first neuron-level modularization framework for general-purpose DNNs, including Transformers. Its key contributions are: (1) scalable, neuron-granularity module decomposition; and (2) a module-aware contrastive learning framework with a composite loss function that jointly optimizes module discriminability and compactness. Extensive experiments across six diverse model architectures demonstrate that NeMo improves module classification accuracy by 1.72% on average, reduces module size by 58.10%, and validates practical deployment efficacy in open-source projects.

Enhancing module accuracy and size efficiency in MwTImproving modularization for diverse and large-scale DNN modelsReducing DNN training costs through modular reuse

Latest Papers

What's happening recently
View more

A Comprehensive Study of Bugs in Modern Distributed Deep Learning Systems

Dec 23, 2025
XM
Xiaoxue Ma
🏛️ Hong Kong Metropolitan University | Wuhan University of Technology | City University of Hong Kong | University College London

Distributed deep learning (DDL) frameworks suffer from a lack of systematic understanding of defects, hindering robustness and maintainability. Method: We conduct the first large-scale empirical study of 849 real-world issues across DeepSpeed, Megatron-LM, and Colossal-AI, proposing the first defect taxonomy tailored to specialized DDL frameworks—comprising 34 symptom categories, 28 root-cause categories, and 6 repair patterns—and establishing a phase-aware symptom–cause–repair mapping across six execution stages: initialization, communication, computation, memory, fault tolerance, and scheduling. Results: We find that 45.1% of symptoms—and 95% of communication-related configuration issues—are uniquely distributed in nature; over 60% of defects are resolvable via version/dependency management or distributed tuning. Setup failures, memory anomalies, and performance deviations emerge as the top three distributed-specific defect types. We quantify root-cause distributions per stage and distill reusable repair patterns and engineering best practices, directly supporting enhanced framework robustness.

Analyzes bugs in distributed deep learning frameworksExamines setup, memory, and performance anomalies in systemsIdentifies symptoms, causes, and fixes for training issues

This work addresses the challenge that existing testing approaches for deep learning library APIs struggle to uncover functional bugs that manifest consistently across hardware but exhibit behavioral discrepancies across frameworks. The authors propose Xamt, the first systematic effort to construct and validate 676 groups of functionally equivalent APIs—spanning 2,563 APIs across seven mainstream deep learning libraries—and detect inconsistencies via cross-framework differential fuzzing. Key innovations include establishing cross-framework API equivalence, normalizing parameter roles, group-level execution validation, and a variance-guided strategy for generating diverse input types. In experiments, Xamt uncovered 72 reproducible issues, including four crashes and 68 output inconsistencies; 25 have been confirmed by developers, and 23 have already been fixed.

API testingcross-framework differentialdeep learning libraries

This work addresses the challenges of coupling Fortran-based numerical models with Python deep learning frameworks, particularly low data transfer efficiency and integration complexity. To overcome these issues, the authors propose TorchNWP, a compilation library that leverages LibTorch to statically compile PyTorch models into C/C++ interfaces, enabling efficient embedding via mixed Fortran/C/C++ programming. The framework innovatively supports automatic generation of tangent linear and adjoint models of neural networks at the C/C++ layer, abstracting internal model structures and thereby simplifying the implementation of four-dimensional variational data assimilation systems. Additionally, it facilitates deployment on heterogeneous platforms and supports multi-granularity parallelism. The approach has been successfully integrated into operational numerical weather prediction systems such as CMA-GFS and MCV, significantly improving forecast accuracy and computational efficiency in radiation and non-orographic gravity wave drag parameterizations.

coupling flexibilitycross-language compatibilitydata transfer efficiency

Existing neuro-symbolic systems are often confined to specific paradigms, making it difficult to unify logical reasoning with deep learning and resulting in high barriers to entry. This work proposes a general-purpose neuro-symbolic AI backend framework that, for the first time, enables unified compilation and modular composition of multiple neuro-symbolic languages. The framework automatically compiles high-level logical specifications into optimized arithmetic circuits and seamlessly integrates them into PyTorch workflows. By treating logic as composable components and leveraging automatic differentiation alongside modular encapsulation, the approach significantly lowers the usability barrier for practitioners while offering researchers an efficient platform for rapid prototyping. The implementation is publicly available as open-source code.

arithmetic circuitsdeep learninglogic integration

Industrial-scale recommendation and ranking models feature highly complex and continuously evolving architectures, rendering traditional optimization approaches—based on manual intervention or module-level rules—difficult to scale. This work proposes the first extensible and customizable operator-level automatic transformation framework integrated into PyTorch 2.x. By leveraging FX intermediate representation, the PT2 compiler, predefined pattern matching, and a greedy search algorithm, the framework achieves general-purpose model optimizations while strictly preserving computational semantics. Evaluated on real-world industrial recommendation models, the approach delivers up to 63% inference speedup, a 6% reduction in peak memory usage, and over 400 seconds of compilation time savings. The implementation has been open-sourced as part of PyTorch 2.x.

deep learninggraph optimizationmodel transformation

Hot Scholars

MC

Ming-Chang Yang

Associate Professor, Department of Computer Science & Engineering at Chinese University of Hong
Non-Volatile MemoryMemory/Storage SystemsEmbedded SystemsComputer Systems
LP

Lujia Pan

Noah's Ark Lab, Huawei
Anomaly dectionTime seriesRepresentation learning
EN

Eliya Nachmani

Ben-Gurion University; Google Research
Deep LearningSpeechAudioSignal Processing
YC

Yueh-Cheng Liu

Technical University of Munich
3D VisionRoboticsDeep Learning