Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the verbosity, inefficient guidance, and high computational costs associated with free-form reasoning in small multimodal agents by proposing the SSR framework, which reformulates open-ended generation as selection from predefined natural language candidates. Innovatively replacing generative reasoning with a selection mechanism, this approach integrates teacher-forced prefilling with shared-context KV caching to enable parallel scoring, thereby achieving efficient structured decision-making without introducing auxiliary task heads. Experimental results demonstrate that SSR maintains competitive success rates while reducing single-turn inference latency by over 90% and total latency by 28%–54%.
📝 Abstract
Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Agents
Reasoning Efficiency
Inference Cost
Small Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selection-based Structured Reasoning
Multimodal Search Agents
Parallel Scoring
Inference Efficiency
Shared KV Cache
🔎 Similar Papers
No similar papers found.