Score
Design, implement, and evaluate strategies for inserting adapter modules or conditioning components into transformer architectures by selecting specific injection sites (e.g., particular layers, attention or feed‑forward submodules, residual paths, or token positions) and tuning their placement; analyze how these placement choices affect parameter efficiency, computational cost, representational change, and downstream adaptation performance.
To address computational inefficiency caused by fixed-depth computation in Transformers, this paper proposes a lightweight dynamic depth routing framework. First, Router-Tuning fine-tunes only a compact router module, avoiding full-model retraining. Second, MindSkip introduces an attention-score-guided dynamic layer-skipping mechanism to prevent skipping semantically critical layers. Evaluated on standard benchmarks, the method retains 99.8% of the original model’s accuracy while accelerating inference by 21% and substantially reducing both computational FLOPs and memory footprint. The core contribution lies in the decoupled design of router adaptation and attention-aware layer skipping—enabling fine-grained, robust, on-demand computation allocation with minimal training overhead. This approach achieves a favorable trade-off among inference efficiency, accuracy preservation, and deployment practicality.
This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.
Conventional MoE routing mechanisms assume expert selection relies solely on semantic features, overlooking the potential influence of positional information. Method: We conduct attribution analysis, attention and routing visualization, systematic ablation studies, and statistical analysis of expert activation distributions across token positions. Contribution/Results: We empirically demonstrate—across multiple state-of-the-art MoE architectures (e.g., Switch Transformer, GLaM)—that tokens at distinct positions exhibit strong, consistent preferences for specific experts, revealing a stable spatial bias in routing decisions. This phenomenon, termed “position-aware routing,” constitutes the first evidence that positional encoding critically shapes expert assignment in MoE models. Our findings establish a novel, interpretable phenomenological model of structured expert allocation and introduce position–semantics co-modeling as a principled optimization axis for MoE design—enhancing both routing efficiency and generalization capability.
To address the high computational cost of pretraining diffusion Transformers (DiTs), this work introduces “model grafting”—the first method enabling low-cost, editable architectural modifications to pretrained DiTs. Grounded in activation analysis and insights into attention locality, the approach supports fine-grained edits ranging from operator substitution (e.g., Softmax → gated convolutions or linear attention) to structural reorganization (e.g., serial → parallel blocks). The resulting hybrid architecture integrates gated convolutions, local/linear attention, variable-expansion-ratio MLPs, and convolutional MLPs. On DiT-XL/2, grafting consumes <2% of the original pretraining FLOPs while achieving FID scores of 2.38–2.64. Applied to PixArt-Sigma, it yields a 1.43× inference speedup with <2% degradation in GenEval. A depth-halved parallel variant achieves FID = 2.77—surpassing same-parameter baselines.
This work investigates how Transformers dynamically acquire inductive capabilities during in-context learning (ICL), specifically focusing on the role of “inductive heads” in transitioning from local n-gram pattern recognition to modeling long-range dependencies. Method: We combine theoretical approximation analysis, synthetic task training dynamics modeling, attention decomposition, and mixed-objective trajectory tracking across training. Contribution/Results: We formally characterize the generalized inductive head mechanism for the first time, revealing a sharp, non-gradual phase transition—from 4-gram modeling to inductive head emergence—during training. We quantify the layer- and head-specific contributions to long-range dependency capture and demonstrate that inductive heads constitute the core architectural substrate underlying ICL emergence. Our study provides the first full-training-dynamics evidence and an interpretable framework for understanding how large language models dynamically generalize, bridging mechanistic analysis with empirical learning trajectories.
This study addresses the challenges of missing source data and catastrophic forgetting during fine-tuning when transferring pretrained chip placement models across designs. We propose CARVE, a continual adaptation framework that enables secure knowledge reuse through frozen base policies, immutable expert modules, and task credential mechanisms. Furthermore, we introduce a novel local-validation-based expert selection strategy, establishing theoretical bounds for expected performance guarantees under fixed distributions and worst-case sample complexity with bounded loss. Experimental results demonstrate that our method reduces training time by 58.5% while preserving nearly lossless HPWL gains. Notably, on IBM circuits, CARVE achieves an average 5.76% HPWL improvement without requiring any receiver-side training.
This work addresses the inefficiency and performance degradation caused by standard LoRA’s uniform adaptation across heterogeneous components in mixture-of-experts language models, which overlooks the functional distinctions between attention and recurrent modules. The authors propose a component-aware LoRA adaptation strategy that differentially deploys adapters in sequential (Qwen3.5-0.8B) and parallel (Falcon-H1-0.5B) hybrid architectures. Experimental results demonstrate that adapting only the attention pathways achieves superior performance over full fine-tuning with 5–10× fewer parameters. Moreover, adapter application to recurrent components proves detrimental in sequential architectures—causing a 14.8% performance drop—but beneficial in parallel ones, yielding an 8.6% gain. The parallel architecture further exhibits positive cross-task transferability, whereas the sequential variant suffers from catastrophic forgetting.
This study investigates how to select trainable parameters under a fixed parameter budget in Low-Rank Adaptation (LoRA) to maximize model performance. It reveals that the effectiveness of parameter selection strategies is highly dependent on the training paradigm: under supervised fine-tuning, random and gradient-guided selection perform comparably, whereas only gradient-guided selection yields gains in GRPO-based reinforcement learning fine-tuning. Building on this insight, the authors propose an efficient scoring method that leverages gradient structure to rapidly identify critical parameters. This approach requires less than 0.5% of the full training cost and completes selection within ten seconds, consistently improving performance across models ranging from 1.5B to 8B parameters. The identified critical parameters are predominantly concentrated in the value (V), output (O), and down-projection matrices.
This work addresses the limitation of conventional optimization methods that impose uniform manifold constraints across all Transformer modules, disregarding their distinct geometric preferences in weight space. To remedy this, the authors introduce the Manifold Muon optimizer during GPT-2 pretraining, proposing a module-specific geometric optimization strategy by assigning Stiefel manifold constraints to attention layers and DGram manifold constraints to MLP layers. Experimental results demonstrate that this differentiated configuration substantially enhances training stability and model performance. In contrast, uniform DGram constraints—or inappropriate manifold assignments—tend to induce singular value growth in attention weights, leading to softmax saturation and unstable optimization. The study thus reveals an asymmetric geometric preference among Transformer components and advocates for function-aware customization of optimization geometry.
This work uncovers a deep connection between the internal operations of Transformer layers and the classical numerical algorithm known as the power method. By interpreting the projections in self-attention together with layer normalization as a single iteration of the power method, the authors theoretically demonstrate that token representations progressively align with the dominant eigenvector of the product of the value matrix and the output weight matrix during forward propagation. This study establishes the first rigorous analogy between Transformers and the power method, introducing a novel paradigm for steering model outputs by manipulating the direction of this dominant eigenvector. Leveraging linear algebraic analysis and a shared-weight architecture, the work both analytically characterizes and empirically validates this alignment phenomenon, demonstrating the feasibility of targeted interventions to control model behavior.