Dopamine: Brain Modes, Not Brains

📅 2026-02-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited interpretability of existing parameter-efficient fine-tuning methods at the neuron level, which obscures how models reuse or bypass internal computations. Inspired by neuromodulation, the authors propose an activation-space fine-tuning approach that models adaptation as a “mode-switching” mechanism. By freezing the backbone weights and learning per-layer trainable thresholds and gains for neurons—combined with smooth gating during training and a hardening strategy at inference—the method enables conditional computation and neuron-level attribution. Evaluated on the rotated MNIST task, it introduces only a few hundred parameters per layer, significantly outperforming frozen baselines while achieving partial activation sparsity and offering clear interpretability of individual neuron activations.

Technology Category

Machine Learning: Mixture of Experts (MoE)Cognitive Modeling & Cognitive Systems: Neural Spike CodingNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Parameter-efficient fine-tuning (PEFT) methods such as \lora{} adapt large pretrained models by adding small weight-space updates. While effective, weight deltas are hard to interpret mechanistically, and they do not directly expose \emph{which} internal computations are reused versus bypassed for a new task. We explore an alternative view inspired by neuromodulation: adaptation as a change in \emph{mode} -- selecting and rescaling existing computations -- rather than rewriting the underlying weights. We propose \methodname{}, a simple activation-space PEFT technique that freezes base weights and learns per-neuron \emph{thresholds} and \emph{gains}. During training, a smooth gate decides whether a neuron's activation participates; at inference the gate can be hardened to yield explicit conditional computation and neuron-level attributions. As a proof of concept, we study ``mode specialization''on MNIST (0$^\circ$) versus rotated MNIST (45$^\circ$). We pretrain a small MLP on a 50/50 mixture (foundation), freeze its weights, and then specialize to the rotated mode using \methodname{}. Across seeds, \methodname{} improves rotated accuracy over the frozen baseline while using only a few hundred trainable parameters per layer, and exhibits partial activation sparsity (a minority of units strongly active). Compared to \lora{}, \methodname{} trades some accuracy for substantially fewer trainable parameters and a more interpretable ``which-neurons-fire''mechanism. We discuss limitations, including reduced expressivity when the frozen base lacks features needed for the target mode.
Problem

Research questions and friction points this paper is trying to address.

parameter-efficient fine-tuning
model interpretability
neuromodulation
conditional computation
activation sparsity
Innovation

Methods, ideas, or system contributions that make the work stand out.

parameter-efficient fine-tuning
neuromodulation
activation sparsity
conditional computation
interpretable adaptation
🔎 Similar Papers
No similar papers found.