🤖 AI Summary
This work addresses three key limitations of artificial neural networks: the absence of biologically inspired internal neuronal states, selective inter-neuronal communication, and self-organizing topological structure. To this end, we propose Intelligent Neural Networks (INNs), wherein neurons are modeled as first-class entities endowed with memory and online learning capabilities, and rigid layering is replaced by a fully connected graph topology. We introduce two core innovations: (i) a selective state-space model enabling neuron-specific state evolution, and (ii) an attention-guided routing mechanism that facilitates autonomous neuron activation and dynamic, context-aware communication. These design choices significantly improve training stability and model interpretability. On the Text8 benchmark, INNs achieve 1.705 bits per character (BPC), outperforming standard Transformers and matching optimized LSTMs. Notably, a Mamba-based baseline with comparable parameter count fails to converge, empirically validating the critical role of the graph-structured topology in stabilizing training.
📝 Abstract
Biological neurons exhibit remarkable intelligence: they maintain internal states, communicate selectively with other neurons, and self-organize into complex graphs rather than rigid hierarchical layers. What if artificial intelligence could emerge from similarly intelligent computational units? We introduce Intelligent Neural Networks (INN), a paradigm shift where neurons are first-class entities with internal memory and learned communication patterns, organized in complete graphs rather than sequential layers.
Each Intelligent Neuron combines selective state-space dynamics (knowing when to activate) with attention-based routing (knowing to whom to send signals), enabling emergent computation through graph-structured interactions. On the standard Text8 character modeling benchmark, INN achieves 1.705 Bit-Per-Character (BPC), significantly outperforming a comparable Transformer (2.055 BPC) and matching a highly optimized LSTM baseline. Crucially, a parameter-matched baseline of stacked Mamba blocks fails to converge (>3.4 BPC) under the same training protocol, demonstrating that INN's graph topology provides essential training stability. Ablation studies confirm this: removing inter-neuron communication degrades performance or leads to instability, proving the value of learned neural routing.
This work demonstrates that neuron-centric design with graph organization is not merely bio-inspired -- it is computationally effective, opening new directions for modular, interpretable, and scalable neural architectures.