Score
Design computational modules that implement learned gating mechanisms to selectively route, filter, and fuse complementary information across processing streams. Build and parameterize gates that are learned jointly with the module to control cross‑stream interactions (for example, selective fusion of parallel streams such as real and imaginary components) and evaluate their effect on information flow.
This work addresses the lack of a unified computational interpretation for neural policy gating mechanisms. We propose GateMod, a theoretically grounded gating framework that couples task structure with neural circuit dynamics via the principle of free-energy minimization. GateMod comprises two core components: GateFlow—a continuous-time energy-flow model—and GateNet—a soft-competitive recurrent network—enabling emergent gating for skill composition and behavioral planning. We formally prove GateMod’s global exponential convergence and robustness under perturbations. Empirically, GateMod achieves significant performance gains over state-of-the-art methods in multi-agent cooperative tasks and human multi-armed bandit experiments. Crucially, it provides the first quantitative demonstration of how task structure modulates gating behavior through neural energy dynamics. By offering a computationally precise and empirically testable account, GateMod establishes a principled theoretical foundation for understanding strategy selection in prefrontal–basal ganglia circuits.
Current spiking neural networks (SNNs) predominantly rely on a single synaptic plasticity mechanism, limiting representational capacity and robustness and failing to emulate the brain’s multi-mechanism collaborative learning. To address this, we propose the first brain-inspired multi-plasticity co-training framework, unifying spike-timing-dependent plasticity (STDP), heterosynaptic STDP (Hetero-STDP), and synaptic homeostasis within a single model. Crucially, we introduce an adaptive mechanism-weight allocation strategy that preserves the intrinsic dynamics of each plasticity rule while enabling dynamic, synergistic optimization. The framework supports end-to-end learning on both static image and dynamic event-camera data. Extensive experiments across multiple benchmark datasets demonstrate substantial improvements in accuracy and generalization over state-of-the-art SNN methods. These results validate the critical role of multi-mechanism collaboration in enhancing SNN expressivity and robustness, establishing a general-purpose, biologically grounded training paradigm for efficient, brain-like SNNs.
Traditional leaky integrate-and-fire (LIF) neuron models neglect the dynamic conductance mechanisms inherent in biological neurons, limiting the robustness and computational capability of spiking neural networks (SNNs) under noise perturbations and temporal variations. To address this, we propose the Dynamic Gating Neuron (DGN) model, which— for the first time—formulates dynamic conductance as a biologically interpretable, adaptive gating mechanism to enable selective information flow regulation and perturbation suppression. The DGN model ensures theoretical stability while maintaining neuroscientific plausibility and supports end-to-end training via backpropagation through time. Evaluated on standard temporal benchmarks—including TIDIGITS and the Spiking Heidelberg Digits (SHD) dataset—DGN consistently outperforms LIF across multiple dimensions: noise robustness, temporal precision, and generalization performance. This work establishes a novel paradigm for designing SNNs that are both highly robust and biologically grounded.
This work addresses the prevalent low-frequency bias in lightweight image classification models by systematically analyzing the impact of gating mechanisms on neural network training dynamics from a frequency-domain perspective. We establish, for the first time, a theoretical frequency-domain interpretation of gating operations—specifically, the coupled element-wise multiplication and nonlinear activation—revealing their collaborative modulation of multi-frequency components. Guided by this analysis, we propose GmNet, a lightweight architecture that minimizes low-frequency bias via a frequency-sensitive information flow control structure, overcoming the empirical limitations of conventional gating designs. Leveraging convolution theorem-based frequency-domain insights for principled model design, GmNet achieves superior accuracy and inference efficiency with fewer parameters on benchmarks including ImageNet, significantly outperforming state-of-the-art lightweight models such as MobileNetV3 and EfficientNet-Lite.
This work addresses the fundamental challenge that conventional neural networks cannot safely learn and infer concurrently, as parameter updates during inference often lead to unstable or even undefined outputs. To resolve this, the authors propose the DynamicGate MLP architecture, which decouples gating (routing) parameters from prediction (representation) parameters. For the first time, they rigorously establish—both structurally and mathematically—the sufficient conditions under which concurrent inference and learning are guaranteed to be safe. The approach ensures that valid model snapshots are maintained even under asynchronous or partial parameter updates, guaranteeing that every inference step yields a well-defined forward pass. This enables the deployment of stable online adaptive systems on edge devices, providing both theoretical assurance and practical foundations for continual learning.
This work addresses the inefficiencies of conventional neural networks in data and energy consumption, as well as the limited availability of effective learning algorithms for spiking neural networks (SNNs). To bridge this gap, the authors propose Spark, a modular SNN framework that constructs end-to-end models by composing simple plasticity-based components, enabling continuous, batch-free learning. Spark integrates seamlessly with traditional machine learning pipelines while incorporating biologically inspired continual learning mechanisms, substantially improving data efficiency and practical applicability. The framework’s effectiveness is demonstrated on the sparse-reward CartPole task, where it successfully learns in a continuous setting, highlighting the promise of modular design for efficient SNN training.
This work proposes a novel computational paradigm for artificial general intelligence grounded in set theory and hyperdimensional computing, addressing the limitations of conventional neural networks that rely on continuous weights and matrix operations and struggle to efficiently emulate biological neural coding. By employing sparse binary representations and subset-based pattern matching, the framework unifies associative memory and symbolic processing without scalar weight adjustments. Instead, it leverages topological plasticity within combinatorially expanded hidden layers, enabling memory capabilities to emerge naturally and seamlessly bridging perceptual and symbolic representations. The resulting system supports constant-time information retrieval and maps directly onto in-memory computing hardware, offering a highly energy-efficient pathway toward artificial general intelligence.
This work investigates the trade-off between model expressivity and generalization performance in Mixture-of-Experts (MoE) architectures under communication constraints. For the first time, rate-distortion theory is introduced into MoE analysis by modeling the gating mechanism as a stochastic channel operating at a finite rate. By integrating mutual information-based generalization bounds with the rate-distortion function \(D(R_g)\), the study establishes a quantitative relationship between the gating communication rate and generalization error. A theoretical upper bound on generalization error is derived and validated through synthetic multi-expert model simulations, which demonstrate that reducing the gating rate, while limiting expressivity, can enhance generalization. Based on these insights, the paper proposes a capacity-aware design principle for MoE systems, offering theoretical guidance for efficient model construction in resource-constrained settings.