🤖 AI Summary
This study aims to clarify whether the performance gains of vision-language models in multi-strategy ensembles stem from sampling noise or genuine specialization. To this end, it proposes a statistical model-based noise stripping method, combines redistribution testing to analyze reliability coverage, and conducts systematic comparative experiments using mixture routing and weight averaging techniques. The findings reveal that no significant specialization differences exist among strategies in most scenarios; substantive routing gains emerge only under strong training recipes and specific threshold conditions. By effectively distinguishing superficial ensemble advantages from genuine specialization, this work provides a rigorous evaluation framework for understanding the underlying causes of routing gains.
📝 Abstract
Policies trained from the same base model can appear to solve different problems. An oracle that chooses the best policy for each problem may therefore appear much stronger than any single policy. Selecting the largest estimated success rate also selects favorable sampling errors. We study this effect through reliability coverage, the fraction of problems whose success probability reaches a chosen threshold. Our first test redistributes stored correctness outcomes across policies within each problem. A second also preserves each policy's total successes, accounting for overall quality differences under a specified statistical model. For five training seeds of a seven-billion-parameter vision-language model, redistribution reproduces 0.096 of an estimated 0.113 oracle gap at threshold 0.10. Neither test finds significant evidence at this threshold. Small advantages remain unresolved. Mixtures, routers, voting, and weight averaging show no detectable improvement over their corresponding single-policy baselines. Training policies on different datasets shows little detectable specialization under light post-training, and no router gain. A stronger recipe does create it, both tests detect it, and a router gain appears only at the high thresholds where the specialists separate. In a control with predictable specialization, a router recovers about half the oracle gap. Coverage bounds explain why even a genuine oracle advantage need not yield a deployment gain. The tests assess apparent specialization from stored responses before investment in routing. Code: https://github.com/KurbanIntelligenceLab/oracle-gaps