Score
Designing continuous-time analog networks and equivalent circuit topologies to implement mathematical functions or algorithms in hardware while meeting area, power, accuracy and stability constraints and enabling closed-loop emergent computation.
Large-scale coupled oscillator networks—such as power grids and neuromorphic systems—face severe computational bottlenecks in simulation due to high arithmetic intensity and hardware inflexibility. To address this, we present a 28 nm reconfigurable on-chip oscillator network chip. Our approach introduces a clustered phase-locked loop (PLL) architecture, with each cluster integrating seven programmable oscillators, co-designed with an embedded RISC-V coprocessor to enable runtime reconfiguration of both network topology and complexity—the first such capability demonstrated in hardware. The chip incorporates custom analog coupling circuits and a brain-inspired on-chip interconnect fabric, and has been silicon-verified for simulations involving hundreds of oscillators. Measurements show two orders-of-magnitude lower power consumption compared to pure digital simulation, significantly improving energy efficiency and adaptability for analog computing and dynamic modeling of critical infrastructure. This work establishes a new paradigm for hardware-accelerated analog computation.
This work reveals that impedance-network-based analog computing is fundamentally constrained by dynamic limitations when performing matrix operations, rendering it incapable of resolving patterns evolving faster than a certain speed. To address this, the study introduces, for the first time, the concept of an “analog Courant number” and defines a composite unity-gain bandwidth (CUGBW), establishing a quantitative relationship between CUGBW and circuit coupling topology. This leads to a theoretical upper bound on the output response rate of analog hardware. Through circuit-theoretic analysis and large-scale LTspice simulations—spanning both CMOS and thermionic vacuum-tube architectures—the derived limit is validated across diverse tasks, including one-dimensional heat equation solving, graph-based semi-supervised learning, and graph-regularized regression. The results demonstrate that the maximum normalized pattern speed cannot exceed 2π times the peak CUGBW, providing a critical theoretical foundation for analog accelerator design.
This work addresses the inefficiency of conventional physical neural networks, which treat nonlinear devices merely as scalar weights and struggle to approximate smooth functions effectively. Inspired by the Kolmogorov–Arnold representation theorem, the authors propose a novel architecture that places trainable nonlinear functions along network connections rather than at nodes, thereby transforming each physical link into a learnable computational unit. This design substantially reduces network size and enhances modeling efficiency for smooth functions, such as those encountered in continuous control tasks, while enabling cross-hardware portability. Implemented using analog bandpass filters based on field-programmable analog arrays, the approach demonstrates architectural generality across both CMOS and memristor platforms. High-accuracy performance is achieved in robotic kinematics, continuous control, and photovoltaic maximum power point tracking with drastically fewer nodes and connections—e.g., only 35,000 connections—than conventional multilayer perceptrons, consuming approximately 30 microwatts when deployed on CMOS hardware.
This work addresses the challenge of efficiently implementing the learnable nonlinear edge functions in Kolmogorov–Arnold Networks (KANs) in hardware by proposing an analog KAN architecture based on a Reconfigurable Nonlinear Processing Unit (RNPU). Leveraging multi-terminal nanosilicon devices that natively support programmable nonlinear transformations, the design employs the RNPU as its fundamental computational element, integrated with analog mixed-signal interfaces to achieve high parameter efficiency, low power consumption, and minimal silicon area for edge neural network deployment. Experimental results demonstrate that, at comparable approximation error levels, the proposed architecture reduces energy consumption by two to three orders of magnitude and chip area by approximately one order of magnitude relative to digital fixed-point MLP implementations, achieving a single-inference energy cost of merely 250 pJ with a latency of about 600 ns.
Analog circuits composed of voltage sources, linear resistors, ideal diodes, and voltage-controlled voltage sources (VCVSs) lack a rigorous theoretical foundation for universal function approximation. Method: We propose a structured mapping that establishes an exact equivalence between ReLU neural networks and such nonlinear analog circuits under ideal-component assumptions. Contribution/Results: We provide the first rigorous proof that these circuits can approximate any continuous function to arbitrary precision, thereby establishing their computational completeness. This result furnishes a solid theoretical basis for self-learning analog neural networks. Furthermore, the proposed circuit architecture is compatible with Equilibrium Propagation—a biologically plausible training framework—and supports end-to-end differentiable analog hardware implementation. By bridging the gap between analog circuit theory and neural computation, our work overcomes the longstanding limitation that analog circuit representational capacity lacks formal theoretical guarantees.
This work addresses the urgent need for efficient solutions to differential and matrix equations in artificial intelligence and scientific computing by transcending the energy-efficiency and speed limitations of conventional digital computation. It pioneers a unified framework that integrates both classes of equations within a modern analog computing paradigm. Leveraging hardware platforms such as analog CMOS circuits and memristor crossbar arrays, the study systematically constructs a computational primitive centered on matrix-vector multiplication, thereby uncovering intrinsic connections among differential equation solvers, matrix equation solvers, and in-memory computing. The research highlights the superior energy efficiency and parallelism offered by memristor arrays while rigorously examining critical challenges including numerical precision and scalability, ultimately establishing analog computing as a promising enabler for next-generation high-performance computing.
This work addresses the challenges of efficiently implementing complex multivariate functions in flexible electronics, which are constrained by circuit density, power consumption, and non-idealities. For the first time, the Kolmogorov–Arnold representation theorem is introduced into flexible analog circuits, leading to the proposal of Analog Kolmogorov–Arnold Networks (AKANs). A co-optimization framework integrating circuit-level error modeling, spline parameter regularization, and hardware-software joint pruning enables significant reductions in hardware overhead while maintaining or even improving approximation accuracy. Experimental results across multiple benchmarks demonstrate that the proposed approach achieves average savings of approximately 30% in both area and power consumption, with peak reductions reaching 55% and 50%, respectively. The method exhibits strong robustness, generality, and energy efficiency, offering a promising pathway for resource-constrained flexible electronic systems.
This work addresses the challenge of efficiently solving high-precision, large-scale MIMO nonlinear optimization problems using conventional in-memory computing approaches. The authors propose a continuous-time, nonlinear closed-loop in-memory computing architecture that embeds the bounded-constraint zero-forcing decoding problem into a nonlinear feedback dynamical system composed of memristor arrays and current-limiting operational amplifiers, enabling direct solution via physical evolution. This represents the first extension of closed-loop in-memory computing from steady-state linear algebra to continuous-time nonlinear optimization. A mixed-precision iterative refinement method tailored to this system is introduced, supporting high-order modulation schemes such as 256-QAM. Experimental validation on a fabricated chip demonstrates correct dynamic behavior under hardware non-idealities in a 16×16 MIMO system, achieving scalable performance ranging from ultra-low-power approximate to high-accuracy detection.
Existing analog circuits suffer from noise accumulation due to temporal feedback, hindering their ability to support always-on, ultra-low-power recurrent neural networks (RNNs). This work proposes a hardware-software co-designed Bistable Memory Recurrent Unit (BMRU), whose discrete outputs and hysteresis-based dynamics effectively suppress noise buildup while enabling a one-to-one mapping of parameters to a current-mode analog circuit. Transistor-level simulations in 180 nm CMOS demonstrate excellent alignment between hardware and software behaviors. In an end-to-end keyword spotting task, the RNN core achieves sub-microwatt inference power, with recurrent-stage power consumption scaling linearly with state dimensionality—yielding over 20× improvement in energy efficiency compared to conventional approaches. This represents the first scalable, high-fidelity ultra-low-power analog RNN implementation.
This work addresses the significant accuracy degradation in closed-loop analog matrix computation (AMC) circuits caused by non-idealities such as device programming errors, thermal noise, operational amplifier offset, and interconnect resistance, which existing simulation methods struggle to model accurately and efficiently. To overcome these limitations, the paper introduces SimAMC, the first simulator capable of efficiently and accurately simulating closed-loop AMC circuits incorporating a comprehensive set of non-ideal effects. SimAMC employs an alternating iterative algorithm to precisely model real-valued matrix operations and integrates detailed non-ideality models. Experimental results demonstrate that SimAMC achieves excellent agreement with SPICE-level simulations while offering speedups of several orders of magnitude, thereby substantially accelerating the design and evaluation of AMC circuits.