On Emergent Capabilities and Model Merging

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过模型合并方法探讨了新兴能力的变化,发现合并操作对这些未明确训练的目标行为有三种主要影响。
📝 Abstract
Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities cheaply. We ask what this operation does to emergent capabilities: behaviors an artifact carries that were never an explicit training target. Studying two independent testbeds (activation oracles and emergent-misaligned models) across three model families, we find that the answer is threefold. First, merging preserves an emergent capability that both parents carry: merging two misaligned checkpoints retains most of their broad misalignment across the whole mixing range. Second, merging cannot create an emergent capability that is superadditive in its parents: no weighted merge of two single-task oracles reaches the jointly-trained oracle's auditing ability. Third, when only one parent carries the capability, merging dilutes it faster than the trained capability that accompanies it: the gap is significant in most settings. In short, emergent behaviors of an artifact do not compose the way its trained capability does.
Problem

Research questions and friction points this paper is trying to address.

emergent capabilities
model merging
fine-tuned checkpoints
adapters
weights arithmetic
Innovation

Methods, ideas, or system contributions that make the work stand out.

model merging
emergent capabilities
superadditive
capability dilution
trained capability
🔎 Similar Papers
No similar papers found.