DynamicGate MLP Conditional Computation via Learned Structural Dropout and Input Dependent Gating for Functional Plasticity

📅 2026-03-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of enabling input-dependent conditional computation during inference while simultaneously achieving effective training regularization and computational efficiency. The authors propose DynamicGate-MLP, a framework that unifies Dropout-style regularization with conditional computation through a learnable continuous gating mechanism. During training, expected gating values provide regularization, while at inference time, the Straight-Through Estimator yields discrete execution paths that dynamically activate subnetworks. A compute budget constraint based on expected gate utilization is introduced, and layer-weighted relative MACs are used to evaluate efficiency. Experiments across multiple datasets—including MNIST, CIFAR-10, Tiny-ImageNet, Speech Commands, and PBMC3k—demonstrate that the method significantly reduces computational overhead while maintaining competitive performance.

Technology Category

Machine Learning: Hardware-aware MLNatural Language Processing: Learning & Optimization for NLPComputer Vision: Diffusion Models for Vision

Application Category

Economics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsResponsible Web: Machine-in-the-loop, human agency and autonomySystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Dropout is a representative regularization technique that stochastically deactivates hidden units during training to mitigate overfitting. In contrast, standard inference executes the full network with dense computation, so its goal and mechanism differ from conditional computation, where the executed operations depend on the input. This paper organizes DynamicGate-MLP into a single framework that simultaneously satisfies both the regularization view and the conditional-computation view. Instead of a random mask, the proposed model learns gates that decide whether to use each unit (or block), suppressing unnecessary computation while implementing sample-dependent execution that concentrates computation on the parts needed for each input. To this end, we define continuous gate probabilities and, at inference time, generate a discrete execution mask from them to select an execution path. Training controls the compute budget via a penalty on expected gate usage and uses a Straight-Through Estimator (STE) to optimize the discrete mask. We evaluate DynamicGate-MLP on MNIST, CIFAR-10, Tiny-ImageNet, Speech Commands, and PBMC3k, and compare it with various MLP baselines and MoE-style variants. Compute efficiency is compared under a consistent criterion using gate activation ratios and a layerweighted relative MAC metric, rather than wall-clock latency that depends on hardware and backend kernels.
Problem

Research questions and friction points this paper is trying to address.

conditional computation
regularization
input-dependent gating
compute efficiency
functional plasticity
Innovation

Methods, ideas, or system contributions that make the work stand out.

DynamicGate-MLP
conditional computation
learned gating
structural dropout
Straight-Through Estimator
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yong Il Choi
Sorynorydotcom Co., Ltd./AI Open Research Lab