MergeHEIR: Mitigating Multimodal Hallucinations as the Tax of Model Merging

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue that model merging exacerbates multimodal hallucinations, while existing mitigation strategies tend to compromise expert capabilities. To tackle this, we propose MergeHEIR, a framework that leverages a small calibration set to construct inter-layer null-space projectors, suppressing hallucinations through periodic SVD projection while preserving specialized competencies. Furthermore, this work introduces the novel concept of "merging tax" and establishes theoretical guarantees for minimal distortion and maximal dimensionality. Extensive experiments demonstrate that MergeHEIR significantly reduces hallucination rates across diverse configurations while efficiently retaining inherited expert knowledge, thereby achieving an optimal trade-off between hallucination mitigation and capability preservation.
📝 Abstract
Model merging consolidates task-specialized experts into a single deployable model. However, we show that such capability consolidation incurs a merging tax of increased hallucination: across 8 model-merging methods, every merged checkpoint exhibits a higher hallucination rate than the average of its constituent experts. An intuitive approach is to adapt existing hallucination-mitigation methods to the post-merge model, yet this unconstrained adaptation disrupts inherited capabilities, creating a tension between hallucination mitigation and expertise retention. To tackle this challenge, we introduce MergeHEIR, a post-merge adaptation framework designed to reduce this merging tax while preserving expertise inherited from initial experts. Using small expert-task calibration sets, MergeHEIR constructs layer-wise null-space projectors via SVD from task-specific activations collected from the merged checkpoint, and periodically projects the accumulated post-merge displacement onto the resulting null spaces to preserve inherited expertise. Theoretically, we establish minimum-distortion and maximum-dimensionality guarantees, characterize the threshold-controlled adaptation-retention trade-off, and extend perturbation guarantees beyond finite calibration data. Across 24 paired comparisons spanning three MLLM configurations and 8 model-merging methods, MergeHEIR consistently mitigates hallucination while largely preserving inherited expertise, demonstrating a more favorable hallucination-retention trade-off.
Problem

Research questions and friction points this paper is trying to address.

Model Merging
Multimodal Hallucinations
Merging Tax
Expertise Retention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model Merging
Multimodal Hallucinations
Null-space Projectors
SVD
Post-merge Adaptation
🔎 Similar Papers