hardware-aware training

Designs and implements training and optimization procedures that incorporate circuit- and device-level models (including device nonidealities, variability, quantization, and nonlinearities) into loss functions and update rules so the learned weights are robust when mapped to physical hardware. Builds hardware-in-the-loop workflows and directly optimizes physical device or synaptic parameters to produce inference-ready weights and improve post-fabrication accuracy and robustness.

hardware-awaretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Supervised Learning for Analog and RF Circuit Design: Benchmarks and Comparative Insights

Jan 21, 2025
AM
Asal Mehradfar
🏛️ University of Southern California | University of California, Irvine

Low design efficiency and high complexity in analog/RF circuit design hinder rapid prototyping and optimization. Method: We propose a supervised-learning-based parameter-to-performance direct mapping framework applicable to both homogeneous and heterogeneous circuits—including LNAs, PAs, VCOs, and receivers—and systematically evaluate Transformer, random forest, k-NN, and fully connected networks for cross-circuit generalization. We introduce a relative-error-based normalization scheme for fair performance assessment and develop a scalable error-suppression strategy tailored to heterogeneous circuits. Contributions/Results: We reveal that parameter–performance linearity critically governs model selection; achieve mean relative errors of 0.3% for LNAs and 0.23% for receivers; reduce prediction error by 88% via heterogeneous data augmentation; and demonstrate that Transformers excel in strongly nonlinear regimes, whereas k-NN exhibits superior robustness under moderate linearity.

Circuit DesignMachine LearningSupervised Learning

This work addresses the challenge that hardware-aware training often fails to universally compensate for diverse hardware non-idealities, necessitating a clear delineation of the boundaries within which trainable mitigation is feasible. The authors propose a diagnostic framework that models hardware non-idealities as structured perturbations to forward operators and systematically evaluates their compatibility with gradient-based optimization through three theoretical criteria: expected gradient consistency, bounded gradient variance, and non-degenerate sensitivity. By integrating forward perturbation modeling, rigorous gradient analysis, and controlled experiments across six representative distortion types—including read noise and IR drop—the study rigorously distinguishes between compensable and inherently uncompensable hardware distortions. These findings provide both theoretical grounding and practical guidance for hardware-software co-design, highlighting that certain non-idealities fundamentally require correction at the circuit or architectural level rather than through algorithmic adaptation alone.

AI acceleratorsgradient-based optimizationhardware-aware training

This work addresses the vulnerability of probabilistic circuits to overfitting and poor generalization under data noise, limited samples, or distribution shifts. To mitigate this, the authors propose PeTeR, the first data-free post-training framework that enhances the robustness of pretrained probabilistic circuits to distributional shifts without requiring retraining. PeTeR leverages distributionally robust optimization by modeling worst-case distributions within a Wasserstein ball and introduces a data-agnostic parameter adjustment mechanism grounded in this principle. Empirical evaluations across multiple density estimation benchmarks demonstrate that PeTeR significantly improves model robustness against both random and adversarial perturbations, matching or outperforming existing data-dependent robust learning baselines.

distribution shiftgeneralizationoverfitting

To address the significant degradation in inference accuracy caused by spatially fixed defects—such as ring-, row-, column-, and checkerboard-patterned faults—in ReRAM-based analog neuromorphic circuits, this work proposes a lightweight neural network-based output voltage correction method. Unlike conventional approaches, it requires no prior knowledge of defect types; instead, it learns a correction mapping solely from the circuit’s raw output voltages, enabling generalization to unseen defect configurations. The method is inherently extensible to dynamic degradation and aging-related faults, supporting real-time adaptive learning. Evaluated within a Design-Technology Co-Optimization (DTCO) simulation framework on the MNIST handwritten digit recognition task, the correction network restores inference accuracy from 55% to 90% under defect conditions—a 35-percentage-point improvement. This work establishes a low-overhead, scalable, and highly robust fault-tolerance paradigm for neuromorphic chips targeting edge and IoT applications.

Correcting inference errors from defects in analog neuromorphic circuitsModeling spatial defect types in multi-layer ReRAM arraysRecovering accuracy loss with lightweight neural network correction

In traditional Bayesian optimization, fixed circuit embeddings often fail to align structural representations with figures of merit (FoM), thereby limiting search efficiency. This work proposes TTARO, a novel framework that introduces test-time representation adaptation for the first time in analog circuit topology search. Within the Bayesian optimization loop, TTARO jointly learns nonlinear feature transformations and a Gaussian process surrogate model online, dynamically adjusting circuit embeddings to continuously align with the optimization objective. By integrating pretrained embeddings, deep kernel learning, and Gaussian process regression, the method enables co-evolution of embeddings and the target FoM. Experiments across 40 configurations demonstrate that TTARO reduces the average simple regret AUC by 15.2% compared to conventional Bayesian optimization and outperforms deep kernel learning baselines by 20.7%, with improvements reaching up to 46.7% in certain scenarios.

analog circuit designBayesian optimizationcircuit embedding

Latest Papers

What's happening recently
View more

Deploying neural networks on unconventional hardware requires balancing accuracy, energy consumption, and platform-specific physical non-idealities, yet existing neural architecture search (NAS) methods are often confined to specific hardware platforms, lacking cross-platform generalization and fair comparability. To address this, this work proposes UH-NAS, a hardware-agnostic framework that, for the first time, leverages large language models as evolutionary operators within a co-design pipeline integrating pluggable hardware backends, platform-level energy models, and non-ideality simulators. Experiments demonstrate that UH-NAS discovers more diverse and robust architectures on unconventional substrates such as optical Mach–Zehnder interferometer (MZI) arrays, significantly outperforming both conventional and LLM-driven NAS baselines. Ablation studies further confirm the critical roles of system prompting and hardware-software co-design in achieving these gains.

cross-platform generalizationhardware-aware optimizationneural architecture search

This work addresses the fragmented nature of existing neural network processor design across training, mapping, and manufacturing stages, which hinders joint optimization of performance, cost, and yield under uncertainty. The authors propose a unified framework grounded in monotonic co-design theory, decoupling training, mapping, manufacturing, and resource allocation through a functional-resource interface to enable both independent optimization and global coordination. A key innovation is the explicit introduction of “confidence” as an optimizable resource within the design flow, allowing uncertainty to be formally modeled while guaranteeing that local improvements automatically advance the global Pareto front. The efficacy of the approach is demonstrated through three case studies: reproducing Pareto-optimal solutions in heterogeneous scenarios, validating confidence as a continuously tunable parameter, and achieving global performance gains without requiring hardware reconfiguration.

end-to-end co-designfabrication yieldhardware-software co-design

This study addresses the degradation of AI model performance in semiconductor manufacturing caused by process variations, equipment aging, and raw material shifts. Leveraging five years of real production line data, the work systematically evaluates multiple MLOps retraining strategies for predictive quality and integrates conformal prediction to deliver statistically valid uncertainty quantification. The authors propose an efficient fixed-interval retraining strategy—updating the model every five lots without hyperparameter tuning—that maintains high prediction accuracy under both abrupt process shifts and gradual equipment degradation while substantially reducing computational overhead. By combining normalized residual control limits with conformal prediction intervals, the approach transitions quality assurance from reactive inspection to proactive, reliable forecasting, offering a robust and practical solution for industrial AI deployment.

MLOpsModel driftPredictive quality

Efficiently training high-energy-efficiency analog resistive networks for machine learning under physical hardware locality constraints remains challenging. This work proposes an analytical gradient computation framework grounded in graph theory and Kirchhoff’s laws, establishing a unified generalized equilibrium propagation model that encompasses both equilibrium propagation and coupled learning. For the first time, this approach enables exact gradient-based training without requiring duplicate network copies. The method achieves localized weight updates using only output-layer information and supports selective tuning of a subset of resistors with minimal performance degradation. Numerical simulations confirm its convergence and effectiveness, offering a novel pathway toward hardware-friendly, brain-inspired computing architectures.

analog computingenergy efficiencylocal learning

Hot Scholars

GZ

Georgios Zervakis

Assistant Professor, Computer Engineering & Informatics, University of Patras
Approximate ComputingDesign AutomationDigital DesignMachine Learning
HF

Holger Fröning

Professor a Heidelberg University
Resource-efficient machine learningneural networksGPUHPC
XS

Xiaobo Sharon Hu

University of Notre Dame, Department of Computer Science and Engineering
Power and reliability aware system-level designCircuit and architecture design for beyond-CMOS devicesAlgorithm and hardware
ZY

Zheyu Yan

Assistant Professor (ZJU100 Young Professor) at Zhejiang University
WW

Wujie Wen

Associate Professor, Department of Computer Science, NC State University
Efficient HardwareDesign AutomationSecure and Private AI ComputingMachine Learning