Score
Designs, implements, and evaluates neural components that control or route information flow by producing multiplicative or switch-like gating signals—including adaptive, channel-wise, decision, dynamic, sparse, and temporal gates—to enable conditioned activation, module selection or fallback switching, and efficient sparse computation. Builds gating layers and their runtime implementations, and analyzes their correctness, routing behavior, and computational tradeoffs (e.g., sparsity, latency, and memory) across training and inference.
This work addresses the lack of a unified computational interpretation for neural policy gating mechanisms. We propose GateMod, a theoretically grounded gating framework that couples task structure with neural circuit dynamics via the principle of free-energy minimization. GateMod comprises two core components: GateFlow—a continuous-time energy-flow model—and GateNet—a soft-competitive recurrent network—enabling emergent gating for skill composition and behavioral planning. We formally prove GateMod’s global exponential convergence and robustness under perturbations. Empirically, GateMod achieves significant performance gains over state-of-the-art methods in multi-agent cooperative tasks and human multi-armed bandit experiments. Crucially, it provides the first quantitative demonstration of how task structure modulates gating behavior through neural energy dynamics. By offering a computationally precise and empirically testable account, GateMod establishes a principled theoretical foundation for understanding strategy selection in prefrontal–basal ganglia circuits.
Traditional leaky integrate-and-fire (LIF) neuron models neglect the dynamic conductance mechanisms inherent in biological neurons, limiting the robustness and computational capability of spiking neural networks (SNNs) under noise perturbations and temporal variations. To address this, we propose the Dynamic Gating Neuron (DGN) model, which— for the first time—formulates dynamic conductance as a biologically interpretable, adaptive gating mechanism to enable selective information flow regulation and perturbation suppression. The DGN model ensures theoretical stability while maintaining neuroscientific plausibility and supports end-to-end training via backpropagation through time. Evaluated on standard temporal benchmarks—including TIDIGITS and the Spiking Heidelberg Digits (SHD) dataset—DGN consistently outperforms LIF across multiple dimensions: noise robustness, temporal precision, and generalization performance. This work establishes a novel paradigm for designing SNNs that are both highly robust and biologically grounded.
This work addresses the prevalent low-frequency bias in lightweight image classification models by systematically analyzing the impact of gating mechanisms on neural network training dynamics from a frequency-domain perspective. We establish, for the first time, a theoretical frequency-domain interpretation of gating operations—specifically, the coupled element-wise multiplication and nonlinear activation—revealing their collaborative modulation of multi-frequency components. Guided by this analysis, we propose GmNet, a lightweight architecture that minimizes low-frequency bias via a frequency-sensitive information flow control structure, overcoming the empirical limitations of conventional gating designs. Leveraging convolution theorem-based frequency-domain insights for principled model design, GmNet achieves superior accuracy and inference efficiency with fewer parameters on benchmarks including ImageNet, significantly outperforming state-of-the-art lightweight models such as MobileNetV3 and EfficientNet-Lite.
To address the challenge of simultaneously achieving biological plausibility, hardware efficiency, and competitive performance in spiking neural networks (SNNs) for general supervised classification, this paper proposes a columnar hierarchical SNN architecture tailored for classification. It employs intra-class-difference-driven columnar organization—each column represents a discriminative subcategory—and adopts an all-spiking signal flow with functionally specialized neurons. A biologically grounded learning mechanism is introduced, integrating local anti-Hebbian plasticity with dopamine neuromodulation to replace backpropagation entirely. The method unifies model-driven reinforcement learning with a state-proximity evaluation framework, enabling end-to-end training directly in the spike domain. Experiments demonstrate that the architecture achieves high accuracy and strong generalization across multiple benchmark classification tasks, while exhibiting exceptional compatibility with low-power neuromorphic hardware. This work establishes a novel paradigm for practical, deployable SNNs.
研究探讨了在稀疏注意力机制中,学习到的门控与随机门控的效果差异,通过控制实验揭示模型表示与施加掩码的共适应导致学习门控优势有限。
Existing spiking neural networks struggle to simultaneously achieve trainability, dynamic diversity, and activity sparsity in temporal regression tasks, often suffering from discretization errors and noise sensitivity. This work proposes a differentiable spiking neuron model based on multi-timescale conductances, which integrates fast, slow, and ultraslow conductance dynamics to jointly modulate current–voltage characteristics. Within a unified architecture, the model naturally supports diverse firing patterns—including tonic, phasic, and burst spiking—while enabling end-to-end gradient-based learning without surrogate gradients. The design is both circuit-realizable and exhibits controllable excitability. Evaluated on the Mackey–Glass time-series prediction task, the proposed model significantly outperforms standard LIF and AdLIF neurons, achieving higher prediction accuracy with substantially lower spike activity density.
This work addresses the lack of a unified formulation among conventional activation functions and their hardware deployment limitations imposed by ADC/DAC bottlenecks. The authors propose Threshold Gating (TG) as a universal primitive for neural nonlinearities, employing an input-conditioned branch gating mechanism that subsumes major activation functions as special cases. Building on this unification, they introduce the “Minimal Branch Theorem,” enabling lossless activation function conversion and seamless model transfer across diverse architectures—such as CNNs, Transformers, and RNNs—without retraining. The framework supports both training from scratch and post-hoc adaptation, offering benefits in model compression, performance enhancement, and training acceleration. Furthermore, it maps efficiently onto analog in-memory computing architectures, substantially reducing power consumption and area overhead.
This work addresses the challenge of enabling input-dependent conditional computation during inference while simultaneously achieving effective training regularization and computational efficiency. The authors propose DynamicGate-MLP, a framework that unifies Dropout-style regularization with conditional computation through a learnable continuous gating mechanism. During training, expected gating values provide regularization, while at inference time, the Straight-Through Estimator yields discrete execution paths that dynamically activate subnetworks. A compute budget constraint based on expected gate utilization is introduced, and layer-weighted relative MACs are used to evaluate efficiency. Experiments across multiple datasets—including MNIST, CIFAR-10, Tiny-ImageNet, Speech Commands, and PBMC3k—demonstrate that the method significantly reduces computational overhead while maintaining competitive performance.
研究提出Gated Spike Axial Propagation机制,解决脉冲变换器中基于稀疏二进制表示的长程通信问题,通过传播-选择过程实现结构化信息传递。
This study addresses the lack of general design principles linking neural network architecture to computational capacity. By systematically evaluating the computational performance of recurrent neural networks with diverse connectivity patterns on Boolean function tasks through large-scale sampling, the work reveals— for the first time—that local 2-cycles and 3-cycles are critical structural motifs for enhancing computational power. It further demonstrates that introducing a small number of sparse connections and biologically inspired interneuron-like units significantly boosts the performance of large-scale networks. The authors construct a comprehensive performance map linking small-network architectures to Boolean function realization, showing that networks containing short cycles achieve optimal performance. Moreover, network performance can be accurately predicted from structural statistics, offering a theoretical foundation and biologically inspired guidance for future neural architecture design.