🤖 AI Summary
This study addresses the limitation of existing model merging methods in disentangling pre-training variations from reinforcement learning (RL) updates, which renders full-parameter transfer ineffective. To overcome this, we propose Selective-RL, a framework that isolates RL parameter updates from a training-stage perspective for the first time. By leveraging principal component analysis to extract dominant directions combined with magnitude preservation techniques, these updates are precisely transferred to the language modules of target vision-language models. We demonstrate that dominant directions exhibit superior cross-model transferability compared to full-parameter updates. Extensive evaluations across three model families and five benchmarks show that our method outperforms full-parameter interpolation in 12 out of 15 cases, achieving a notable improvement of 8.55 percentage points for Qwen on MathVision.
📝 Abstract
Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL). Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update. Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise directions with magnitude preservation, and transfers them to the language modules of a VLM. Across three model families and five visual-reasoning benchmarks, Selective-RL improves full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient. Matched controls show that update magnitude or arbitrary low rank alone does not reproduce these gains. These results highlight a distinction between what is acquired during post-training and what remains transferable across models, providing a training-stage perspective on cross-model capability transfer. Code is available at https://anonymous.4open.science/r/selective-rl.