probabilistic gating

Designs, implements, and evaluates mechanisms that probabilistically attenuate or route signals, criteria, or model outputs—e.g., differentiable or stochastic gates, soft gating functions, and soft suppression—so gates can be trained end-to-end instead of using hard on/off decisions. Work includes choosing gate parameterizations, integrating gating into objective and reward computations, deriving/estimating gradients and variance, and measuring effects on downstream utility, leakage, and stability.

probabilisticgating

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Easing Optimization Paths: a Circuit Perspective

Jan 04, 2025
AO
Ambroise Odonnat
🏛️ Noah's Ark Lab | Inria | FAIR | Meta AI

This study addresses the high computational cost and poor safety controllability in training ultra-large AI models. Methodologically, it introduces a novel optimization paradigm grounded in mechanistic interpretability, pioneering the application of “circuit analysis” to model gradient descent trajectories. It structures the parameter space into functional subnetworks and designs a progressive curriculum learning strategy to dynamically regulate optimization paths within a controlled environment. Key contributions include: (1) establishing a formal mapping between gradient flow dynamics and circuit-like structural representations, enabling interpretable modeling of optimization behavior; and (2) leveraging structural priors to guide curriculum design, significantly accelerating convergence while suppressing the emergence of harmful behaviors. Experiments across multiple benchmark tasks demonstrate over 30% reduction in training cost alongside improved behavioral controllability, offering a principled pathway toward efficient and safe large-model training.

Large-scale AI systemsOptimizationSafe learning

Neural Policy Composition from Free Energy Minimization

Dec 04, 2025
FR
Francesca Rossi
🏛️ Scuola Superiore Meridionale | ETH | UC Santa Barbara | University of Salerno

This work addresses the lack of a unified computational interpretation for neural policy gating mechanisms. We propose GateMod, a theoretically grounded gating framework that couples task structure with neural circuit dynamics via the principle of free-energy minimization. GateMod comprises two core components: GateFlow—a continuous-time energy-flow model—and GateNet—a soft-competitive recurrent network—enabling emergent gating for skill composition and behavioral planning. We formally prove GateMod’s global exponential convergence and robustness under perturbations. Empirically, GateMod achieves significant performance gains over state-of-the-art methods in multi-agent cooperative tasks and human multi-armed bandit experiments. Crucially, it provides the first quantitative demonstration of how task structure modulates gating behavior through neural energy dynamics. By offering a computationally precise and empirically testable account, GateMod establishes a principled theoretical foundation for understanding strategy selection in prefrontal–basal ganglia circuits.

Derives a normative framework for policy gating via free energy minimizationDevelops a computational model linking gating to task structure and neural circuitsProvides interpretable explanations of gating in multi-agent systems and decision-making

Deep Learning Meets Mechanism Design: Key Results and Some Novel Applications

Jan 11, 2024
VU
V. Udaya Sankar
🏛️ SRM University A.P. | Carnegie Mellon University | Indian Institute of Science

Mechanism design faces fundamental trade-offs among incentive compatibility, individual rationality, social welfare, and revenue maximization—constraints often mutually incompatible under classical analytical approaches. This paper introduces the first systematic integration of deep learning with classical mechanism design theory, proposing an end-to-end learnable mechanism framework: mechanisms are parameterized by neural networks; game-theoretic equilibrium constraints are explicitly embedded into the model architecture; and a multi-objective customized loss function is optimized jointly via backpropagation. Our approach overcomes the limitations of analytical construction, achieving approximate Pareto-optimality even for theoretically infeasible combinations of desiderata. We validate the framework on three real-world applications—vehicular energy management, mobile network resource allocation, and agricultural collective procurement auctions—demonstrating significant improvements in the balance between social welfare and platform revenue.

Deep learning approximates theoretically incompatible mechanism design propertiesIt addresses real-world applications like energy and resource managementThe approach is demonstrated through case studies in various networks

Quadratic Gating Functions in Mixture of Experts: A Statistical Insight

Oct 15, 2024
PA
Pedram Akbarian
🏛️ The University of Texas at Austin | Johns Hopkins University

This work addresses two key limitations in mixture-of-experts (MoE) models: the lack of theoretical connection between MoE routing and self-attention, and the low sample efficiency of linear gating. We propose quadratic gating—replacing conventional linear routing with a quadratic function—and establish, for the first time, its rigorous equivalence to self-attention. Leveraging this equivalence, we derive principled design criteria for optimal quadratic gating and expert functions, leading to a novel high-performance attention mechanism. Theoretically, via statistical learning analysis, we prove that quadratic gating substantially enhances the expressivity and parameter/sample efficiency of expert selection. Empirically, our MoE variant outperforms linear-gating baselines across multiple tasks; the new attention mechanism surpasses state-of-the-art methods—including FlashAttention and Multi-Head Attention—while exhibiting strong alignment between theoretical predictions and empirical results. The framework thus achieves both interpretability and practical efficacy.

Analyzes convergence of MoE models with quadratic gating functionsEstablishes connection between MoE and self-attention mechanismsProposes active-attention mechanism to enhance self-attention performance

Latest Papers

What's happening recently
View more

This work proposes the first cost-aware routing framework for supervised fine-tuning data acquisition that integrates statistical gating with an adversarial adjudication mechanism to efficiently identify high-value corpora while avoiding costly misacquisitions. The approach evaluates candidate samples along three axes—diversity, utility, and redundancy—using low-cost statistical estimates for initial filtering and triggering a multi-agent debate between proponent and opponent advocates only when confidence is insufficient. Evaluated through quality assessments with confidence intervals and controlled synthetic benchmarks, the system achieves 0.90 accuracy and 0.83 F₁ score across twelve datasets at a unit cost of just $0.017, substantially outperforming always-verify strategies. Moreover, it provides the first quantitative evidence of stance bias (52% stance reversal) and oppositional advantage (80% win rate) in LLM-based adjudication.

cost-aware decisiondata procurementquality assessment

This work addresses the abstraction gap between differentiable programming and emerging probabilistic hardware by introducing Parameterized Stochastic Circuits (PSCs) as a gate-level intermediate representation that closely mirrors native hardware operations. PSCs unify explicit binary, categorical, and continuous signals with localized stochastic kernels. Building on this foundation, the authors develop torx, an open-source JAX framework that enables, for the first time, direct alignment between differentiable stochastic computation and probabilistic hardware, substantially reducing mapping overhead. Experiments on the X0 subthreshold CMOS probabilistic bit chip demonstrate that PSCs efficiently harness physical randomness in tasks such as graph random walks, discrete diffusion, and stochastic graph networks, achieving results consistent with pseudorandom software baselines while confirming the approach’s validity and hardware energy efficiency.

energy efficiencyhardware-native operationsprobabilistic hardware

This work addresses the high computational overhead of conventional neural networks that hinders real-time EEG classification on edge devices. It introduces, for the first time, differentiable logic gate networks (Diff-Logic) to EEG classification, compiling them into pure Boolean circuits and leveraging native CPU bitwise operations for hardware-efficient inference. The proposed approach matches or surpasses the performance of multilayer perceptrons (MLPs) while drastically reducing latency and model size. In dementia screening, it achieves a Macro F1 score of 80.2%, outperforming MLPs by 6.8%. For emotion recognition, it reduces model size by 14× and latency by 2.3×, with up to a 2.9× speedup in inference at the largest scale and near-constant inference time—making it highly suitable for resource-constrained edge deployment.

edge devicesEEG classificationlow-latency

This work addresses the lack of a unified structural and loss-based interpretation for existing activation functions—such as GELU, ReLU, and SiLU—which hinders the principled design and understanding of novel variants. The authors propose a new perspective grounded in Gaussian-complementary first-order loss, interpreting GELU as a hard linear gating signal with a Gaussian-distributed random threshold. This insight leads to a generalized “threshold–transmission” family of activation functions that explicitly incorporates an adjustable threshold width parameter, accommodating both fixed and learnable uniform-threshold variants. Through probabilistic threshold gating, piecewise polynomial construction, and controlled empirical validation, the proposed activations consistently match or outperform established alternatives across compact vision and language models, while efficiently leveraging a limited transition region.

activation functionfirst-order lossGELU

Hot Scholars

SY

Shihao Yang

Assistant Professor, School of Industrial & Systems Engineering, Georgia Institute of Technology
Digital Disease DetectionElectronic Health RecordsMarkov Chain Monte CarloDynamic System Inference
NH

Nhat Ho

Assistant Professor at University of Texas, Austin
Machine LearningBayesian StatisticsOptimizationOptimal Transport
JD

Jiankang Deng

Imperial College London
Computer VisionMachine Learning
YG

Yeyun Gong

Microsoft Research Asia
Natural Language GenerationQuestion AnsweringPre-training