ReSight-SMC: Two-Stage Power Sampling via Island SMC with Visual Scouts

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses trajectory collapse and insufficient visual exploration induced by power sampling in large vision-language models (LVLMs) by proposing a verifier-free, two-stage sampling framework. Methodologically, it introduces an ancestor-isolated island Sequential Monte Carlo (SMC) mechanism to preserve reasoning trajectory independence, alongside a prefix-conditioned visual scout that optimizes attention allocation over image regions. The framework further integrates sequence-level power sampling, exact importance weight correction, and canonical-answer-based second-stage aggregation. Experiments demonstrate that the proposed method consistently outperforms Power-SMC across four LVLM backbones and five benchmarks, achieving performance comparable to reinforcement learning-based models without requiring post-training.
📝 Abstract
Power sampling has emerged as a training-free approach to LLM reasoning, eliciting capabilities comparable to reinforcement learning by sharpening the model distribution over complete responses. Despite this success, power sampling remains underexplored in large vision-language models (LVLMs). We transfer Power-SMC to LVLM decoding by defining a sequence-power target conditioned on both the image and the prompt. This direct transfer provides a strong training-free baseline, but leaves two aspects of finite-particle multimodal inference unaddressed. At the particle level, global resampling can collapse genealogies, while particle-based power sampling does not diversify trajectories through distinct visual cues in multimodal decoding, limiting exploration under a finite particle budget. At the answer level, sequence-level sharpening makes distinct reasoning trajectories compete even when they support the same answer. We introduce ReSight-SMC, a verifier-free two-stage power sampler for LVLM inference. Its first stage uses ancestry-isolated SMC islands to preserve independent trajectory families and routes a bounded set of prefix-conditioned visual scouts to prefix-relevant image regions while discouraging redundant overlap. Each scout temporarily increases attention to the image tokens and emphasizes its routed region. Exact importance correction preserves the base LVLM sequence-power target. The second stage aggregates terminal importance mass by canonical answer, powers the answer marginal, and samples an answer together with a supporting trajectory. Across four LVLM backbones and five benchmarks, ReSight-SMC achieves stronger aggregate performance than Power-SMC over both the reasoning and perception benchmark groups. Without post-training, it remains competitive in aggregate with backbone-matched models trained using reinforcement learning.
Problem

Research questions and friction points this paper is trying to address.

power sampling
large vision-language models
sequential Monte Carlo
multimodal reasoning
particle collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Power Sampling
Sequential Monte Carlo
Large Vision-Language Models
Visual Scouts
Training-free Inference
🔎 Similar Papers
No similar papers found.