Score
Designs and implements mechanisms that dynamically control how intermediate feature representations are selected, weighted, fused, or suppressed within a model—covering selective feature gating, feature fusion gating, and adaptive gating mechanisms. These methods modulate feature responses based on input, context, or auxiliary signals to amplify causally useful traits and reduce the influence of distracting or spurious cues.
This work addresses the prevalent low-frequency bias in lightweight image classification models by systematically analyzing the impact of gating mechanisms on neural network training dynamics from a frequency-domain perspective. We establish, for the first time, a theoretical frequency-domain interpretation of gating operations—specifically, the coupled element-wise multiplication and nonlinear activation—revealing their collaborative modulation of multi-frequency components. Guided by this analysis, we propose GmNet, a lightweight architecture that minimizes low-frequency bias via a frequency-sensitive information flow control structure, overcoming the empirical limitations of conventional gating designs. Leveraging convolution theorem-based frequency-domain insights for principled model design, GmNet achieves superior accuracy and inference efficiency with fewer parameters on benchmarks including ImageNet, significantly outperforming state-of-the-art lightweight models such as MobileNetV3 and EfficientNet-Lite.
This work addresses the limited reconstruction performance in U-Net decoders caused by imprecise fusion of high- and low-level features. To overcome this, the authors propose a novel difference-driven adaptive gating mechanism that, for the first time, leverages the discrepancy between high- and low-level feature streams—rather than their content or correlation—to generate coupled gating maps that precisely modulate the fusion process. Two specific implementations are introduced: Feature Difference-based Gating (FDG), which uses the absolute difference of features for local refinement, and Entropy Difference-based Gating (EDG), which exploits signed entropy differences to enable global optimization. Extensive experiments demonstrate that the proposed approach consistently outperforms existing attention-based fusion strategies across diverse tasks, including medical image segmentation, remote sensing cloud removal, and speech separation, with EDG achieving the best overall performance.
This work addresses the lack of a unified computational interpretation for neural policy gating mechanisms. We propose GateMod, a theoretically grounded gating framework that couples task structure with neural circuit dynamics via the principle of free-energy minimization. GateMod comprises two core components: GateFlow—a continuous-time energy-flow model—and GateNet—a soft-competitive recurrent network—enabling emergent gating for skill composition and behavioral planning. We formally prove GateMod’s global exponential convergence and robustness under perturbations. Empirically, GateMod achieves significant performance gains over state-of-the-art methods in multi-agent cooperative tasks and human multi-armed bandit experiments. Crucially, it provides the first quantitative demonstration of how task structure modulates gating behavior through neural energy dynamics. By offering a computationally precise and empirically testable account, GateMod establishes a principled theoretical foundation for understanding strategy selection in prefrontal–basal ganglia circuits.
This work addresses two key limitations in mixture-of-experts (MoE) models: the lack of theoretical connection between MoE routing and self-attention, and the low sample efficiency of linear gating. We propose quadratic gating—replacing conventional linear routing with a quadratic function—and establish, for the first time, its rigorous equivalence to self-attention. Leveraging this equivalence, we derive principled design criteria for optimal quadratic gating and expert functions, leading to a novel high-performance attention mechanism. Theoretically, via statistical learning analysis, we prove that quadratic gating substantially enhances the expressivity and parameter/sample efficiency of expert selection. Empirically, our MoE variant outperforms linear-gating baselines across multiple tasks; the new attention mechanism surpasses state-of-the-art methods—including FlashAttention and Multi-Head Attention—while exhibiting strong alignment between theoretical predictions and empirical results. The framework thus achieves both interpretability and practical efficacy.
This study investigates the mechanistic degradation and reversibility of language models under toxic data fine-tuning. Toxic fine-tuning induces model corruption, yet its underlying neural mechanisms and potential for recovery remain poorly understood. Method: Leveraging causal tracing and circuit localization—key techniques from mechanistic interpretability—alongside task-specific fine-tuning and clean-data reverse retraining, we conduct controlled ablation and reconstruction experiments. Results: We establish, for the first time, that corruption exhibits *circuit-level specificity*: only critical computational pathways are selectively impaired, while peripheral circuits remain intact. Crucially, we demonstrate *neuroplastic-like recoverability*: clean-data retraining reconstructs original functional mechanisms with >89% restoration fidelity; this recovery generalizes across fine-tuning epochs. Contribution: Our work identifies precise circuit-level localization principles governing corruption and empirically validates the reversibility of mechanistic damage—providing both theoretical foundations and actionable strategies for robust alignment and trustworthy fine-tuning.
This study addresses the issue in large audio-language models where irrelevant audio interferes with textual reasoning and aggregated accuracy obscures pairwise drift phenomena. To tackle this, we propose the ICAP-Gate mechanism, which leverages pairwise drift analysis and mechanism-guided intervention to precisely localize and control architecture-specific late-stage audio pathways. This approach establishes a design principle for selective modality influence control, enabling task-conditioned selective listening. Extensive evaluations across four models and two benchmarks demonstrate that our method significantly reduces both the influence rate and answer flip rate, effectively suppressing interference while preserving ASR performance. Furthermore, the inference latency of ICAP-Gate is substantially lower than that of self-consistency methods.
This study addresses the unclear internal representational mechanisms underlying refusal behavior in activation steering of large language models. By decomposing refusal steering into component-level interventions through sparse subset analysis and residual stream dimension decoupling, this work identifies sparse attention and MLP subsets capable of reproducing the behavioral effect. It provides the first evidence that refusal is assembled by structured, identifiable sparse mechanisms rather than diffusely encoded representations, thereby establishing the privileged basis hypothesis. Experiments demonstrate that retaining only 28–48% of components or 50% of residual stream dimensions preserves 85–101% of the original steering efficacy, revealing a dual sparsity of the signal across both components and dimensions. The implementation code has been made publicly available.
This study addresses the paradox that while invertible transformations preserve information, they alter the distribution of feature contributions to predictions, yielding a fundamental divergence between information preservation and contribution preservation. To resolve this, we decouple intrinsic information from implementation-dependent contributions and propose quantitative transition laws for multi-representation prediction. By integrating cooperative game-theoretic frameworks with Lipschitz utility analysis, we formalize this deficiency and demonstrate that behavioral distance governs the magnitude of its impact, which is further validated through electrocardiogram experiments. This work elucidates the mechanism by which nonlinear recoding degrades predictive accuracy and establishes that exact inverse mappings can fully recover coalition-wise accuracy across all feature subsets. These findings provide a rigorous theoretical foundation for ensuring contribution consistency in representation learning.
This study addresses the challenge in large language model preference alignment where base models already possess target capabilities but fail to express them reliably. To this end, this work proposes a direct hidden state alignment framework that parses preference representations via Residual Competition Maps (RCMs) and employs Causal Activation State Transition (CAST) to intervene in local activations during inference while keeping the base model frozen. By shifting the adaptation space from model weights to hidden states, the approach enables lightweight control without retraining. Remarkably, the proposed method achieves performance comparable to Direct Preference Optimization (DPO) using only 256 to 16,384 parameters. Furthermore, it supports pluggable, multi-domain preference control and functions complementarily with DPO, offering a highly parameter-efficient alternative for aligning large language models.
This study addresses the prevalent conflation of feature manipulation effects with intrinsic computational mechanisms in large language model (LLM) research, wherein behavioral changes are erroneously taken as evidence of actual feature utilization. To resolve this, we propose an empirical contract grounded in natural input values, employing causal intervention techniques—including activation copying, removal, and downstream recovery—to conduct comparative experiments across diverse LLM representations. This framework effectively disentangles manipulation strength from the degree of model reliance. Our findings reveal that no single metric suffices to demonstrate genuine feature utilization and uncover significant disparities in the extent to which different representations can be manipulated versus utilized. Ultimately, this work corrects the cognitive bias of inferring internal mechanisms solely from external intervention outcomes, offering a more rigorous methodological foundation for mechanistic interpretability in LLMs.