🤖 AI Summary
This study addresses the limited candidate diversity in conventional test-time scaling caused by fixed computational paths. We propose a training-free architectural sampling method that reuses decoder blocks of frozen vision-language models, such as Qwen, to generate diverse forward computation paths. By substituting temperature sampling with computational diversity, this approach enhances candidate coverage and self-learning efficacy while keeping model weights frozen and introducing no auxiliary parameters. Experimental results demonstrate that our technique yields an average 6.58% improvement in pass@9 across twelve multimodal benchmarks. Furthermore, it significantly reduces lexical overlap among generated candidates and optimizes performance in label-free reinforcement learning settings.
📝 Abstract
Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.