SMAT: Simple and Efficient Merge-Aware Training

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation following model merging and the prohibitive training costs of existing approaches by proposing a simplified merge-aware training framework from an operation-centric perspective. The method abstracts model merging into scaling, masking, and perturbation operations. By sampling to simulate merged parameters, it jointly optimizes expert and expected losses, enabling efficient single-step forward-backward propagation. Furthermore, periodic scheduling, kernel fusion, and parameter storage switching are incorporated to minimize overhead. Experimental results demonstrate that this framework yields average performance improvements of 1.07–2.16 points across diverse backbone architectures while incurring less than 2% additional training cost.
📝 Abstract
Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise. We further introduce periodic scheduling, kernel fusion, and parameter storage switching to make SMAT efficient, with one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves the mean score across five merging methods by 1.07-2.16 points over the strongest baseline for each backbone, with less than 2% training-time overhead over standard fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

Model Merging
Merge-Aware Training
Expert Training
Training Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Merge-Aware Training
Model Merging
Scale-Mask-Perturb Abstraction
Kernel Fusion
Periodic Scheduling
🔎 Similar Papers