multi-tool agentic reasoning

Designs and builds agentic systems that integrate multiple external tools with vision-language models to perform grounded reasoning over visual and textual inputs. These systems orchestrate tool invocation for search, spatial focusing, temporal localization and layout generation, maintain closed-loop reasoning trajectories, and synthesize evidence‑based, tool‑augmented answers.

multi-toolagenticreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios

Aug 25, 2025
BZ
Bingxi Zhao
🏛️ Beijing Jiaotong University | Lancaster University | Max Planck Institute for Informatics | University of Electronic Science and Technology of China

Existing LLM-based agent reasoning frameworks lack systematic categorization and comparative analysis. Method: This paper presents the first comprehensive survey, proposing a unified taxonomy covering single-agent, tool-augmented, and multi-agent paradigms; formalizing reasoning structures, control flows, and interaction mechanisms via a rigorous descriptive language; and conducting a systematic literature review complemented by cross-domain comparative analysis (research, healthcare, software engineering) and empirical evaluation. Contribution/Results: The study identifies applicability boundaries, performance bottlenecks, and validation methodologies for each paradigm, establishing the first end-to-end mapping from methodological design to real-world application scenarios. It delivers a structured knowledge graph and authoritative reference for theoretical modeling, benchmark development, and engineering deployment of agent reasoning systems.

Analyzing framework-level reasoning across application scenariosProviding systematic taxonomy and evaluation strategiesSurveying diverse LLM-based agentic reasoning frameworks

From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users

Aug 24, 2025
SS
Sadia Sultana Chowa
🏛️ Daffodil International University | United International University | Charles Darwin University

Existing research on large language models (LLMs) as autonomous agents and tool users remains fragmented and limited in architecture design, multi-agent coordination, tool integration, cognitive mechanism modeling, and evaluation frameworks. Method: This survey systematically analyzes 2023–2025 top-tier conference and journal publications using structured literature analysis, integrating prompt engineering and fine-tuning techniques to dissect LLM implementations of core cognitive capabilities—reasoning, planning, and memory. Contribution/Results: We identify three breakthrough directions—verifiable reasoning, self-improvement, and personalized customization—and distill ten concrete future research pathways. Further, we propose a unified evaluation framework covering 68 publicly available datasets, exposing critical gaps in current benchmarks regarding task generalization, dynamic adaptability, and causal attribution capability.

Analyzing architectural designs for single and multi-agent systemsEvaluating benchmarks and datasets for LLM agent performanceExamining LLMs as autonomous agents and tool users

Must-Read Papers

Most classic and influential ideas
View more

ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use

Oct 31, 2025
MD
Mengjie Deng
🏛️ Renmin University of China

To address the inflexible tool invocation and visual context degradation of multimodal large language models (MLLMs) in long-horizon visual question answering (VQA), this paper proposes an agent framework integrating global planning with local perception. The method introduces a dedicated *Perceive* tool that explicitly models the visual perception process, unifying long-range navigation and fine-grained local understanding; it further establishes a hierarchical reasoning architecture enabling iterative, collaborative scheduling of *Search*, *Code*, and *Perceive* tools. This design effectively mitigates visual information decay during multi-step reasoning and significantly enhances cross-modal alignment and tool generalization. Evaluated on four cross-domain VQA benchmarks, the approach achieves an average improvement of +6.69%, demonstrating its effectiveness and robustness for long-horizon, multi-step, and multimodal reasoning tasks.

Addressing visual context degradation in long-horizon VQA tasksEnabling multimodal LLMs to flexibly use external toolsUnifying global planning with local multimodal perception

This work addresses the limitations of existing vision-language reasoning agents, which often lack adaptive judgment regarding task difficulty, leading to erroneous tool invocations on simple tasks and marginal gains on complex ones. To overcome this, the authors propose a reinforcement learning framework grounded in multimodal large language models, featuring a necessity-aware adaptive reward mechanism and a prompt-guided capability expansion strategy. This approach enables agents to invoke tools selectively based on task demands. The method significantly improves both the accuracy of tool invocation timing and the agent’s ability to solve challenging problems, achieving concurrent gains in overall performance and tool utilization efficacy across multiple benchmarks.

agentic visual reasoningMode Adaptivenessmultimodal large language models

PyVision: Agentic Vision with Dynamic Tooling

Jul 10, 2025
SZ
Shitian Zhao
🏛️ Shanghai AI Lab | CUHK | NUS | Rice University

Existing visual reasoning approaches are constrained by predefined workflows and static toolsets, limiting flexibility and interpretability. This paper proposes a multi-round interactive framework enabling multimodal large models to autonomously generate, execute, and iteratively refine Python-based tools tailored to visual tasks. Its core innovation is a dynamic tool generation mechanism: rather than relying on a fixed tool library, the model synthesizes executable code tools on-the-fly according to task requirements, with execution feedback driving iterative refinement across multiple rounds. Integrating multimodal perception, program synthesis, and closed-loop execution, the framework establishes the first evolvable visual tool invocation system. Experiments demonstrate substantial performance gains across multiple benchmarks: +7.8% on V* for GPT-4.1 and +31.1% on VLMsAreBlind-mini for Claude-4.0-Sonnet, validating both effectiveness and generalizability.

Advancing agentic visual reasoning through dynamic tool inventionEnabling MLLMs to dynamically generate and refine Python-based toolsOvercoming limitations of predefined workflows in visual reasoning

This work addresses the limitations of existing tool-integrated reasoning models, which rely on predefined tools, lack self-optimization capabilities, and incur high tool construction costs, thereby struggling with open-ended tasks. To overcome these challenges, the authors propose UCT, a novel framework that introduces the first training-free paradigm for automatic tool construction. UCT extracts reasoning traces from large language models and distills them into reusable tool assets, while incorporating a memory consolidation mechanism to dynamically maintain and update the tool library. This enables agents to autonomously create and refine tools during reasoning, transitioning from mere tool users to tool creators. Evaluated on diverse mathematical and scientific reasoning benchmarks, UCT achieves performance gains of 20.86% and 23.04%, respectively, demonstrating its capacity for continuous self-improvement.

erroneous tool outputsmanual tool constructionopen-ended problems

Vision-language agents are constrained by reliance on human-annotated supervision, while text-based self-evaluation suffers from hallucination and fails to reliably verify multi-step visual reasoning. Method: This paper proposes a tool-augmented self-evolving reasoning framework that establishes a closed-loop of “reasoning–self-evaluation–self-repair,” enabling continuous autonomous optimization without external rewards or human annotations. Crucially, tool invocation is deeply integrated into the reasoning chain to generate structured self-feedback, and fine-grained self-reward signals are designed to drive reinforcement learning. Contribution/Results: Evaluated on geometric problem solving and visual scientific analysis, the model achieves a 12.5% improvement over strong baselines. This work provides the first empirical validation of stable, self-sustained evolution in vision-language agents under zero external supervision.

Addressing evaluation hallucinations in purely text-based self-assessment methodsEnabling continuous self-improvement through tool-integrated reasoning and verificationOvercoming limitations of human-annotated supervision in vision-language agents

Latest Papers

What's happening recently
View more

Existing vision-language models struggle to effectively compose multiple visual tools and dynamically adjust their reasoning based on feedback. This work presents the first systematic approach to modeling the compositionality and adaptability of visual tool usage, introducing a hierarchical synthetic trajectory construction method and a two-stage training paradigm comprising supervised pretraining followed by reinforcement learning. Central to the framework is a multi-tool composition scheduling mechanism that enables flexible and context-aware tool integration. The proposed method achieves state-of-the-art performance among open-source models, attaining 95.8% accuracy on the V* benchmark and 35.3% on VTC-Bench, while demonstrating strong generalization capabilities in more complex tool-rich environments.

adaptive reasoningcompositional reasoningmultimodal agents

Existing vision-language models struggle to recover fine-grained spatial information in tasks requiring active evidence acquisition and multi-step visual interaction. To address this limitation, this work proposes the PERIA agent, which features a lightweight dual-tool architecture: a perception tool that extracts textual, symbolic, and spatial evidence, and an interaction tool that manipulates visual context, tracks trajectories, and verifies relational structures. By integrating supervised trajectory synthesis with a novel Observation-Relaxed Group-in-Group Policy Optimization (OR-GIGPO) strategy, PERIA enables efficient multi-step reasoning. Evaluated across 13 benchmarks on 8 datasets, PERIA-8B substantially outperforms Qwen3-8B by +10.0% in-distribution and +4.4% out-of-distribution, surpassing current state-of-the-art models by 7.0%–14.8%, and achieves performance comparable to Qwen3-VL-235B-A22B-Thinking and GPT-5.

evidence acquisitionspatial reasoningtool-augmented agents

This work investigates the phenomenon of “tool-use collapse” in visual reasoning agents, where tool invocation frequency declines in late-stage training despite concurrent performance gains—particularly in complex tasks such as 3D spatial reasoning and medical visual question answering. The authors propose treating tools as training scaffolds and introduce entropy regularization to enhance exploration diversity in both language generation and tool usage, thereby promoting varied reasoning pathways rather than merely maximizing invocation frequency. Experimental results demonstrate that even with reduced tool usage, increased diversity in reasoning trajectories significantly improves model performance. This finding reveals a nonlinear relationship between tool-use frequency and reasoning capability, challenging conventional frequency-driven tool optimization paradigms and offering a new perspective on effective tool integration in multimodal reasoning systems.

3D spatial reasoningmedical VQAtool-use collapse

This work challenges the common assumption that tool usage inherently enhances multimodal reasoning capabilities by introducing a distinction between tool availability and actual tool contribution. It systematically evaluates two multimodal agents—Thyme and DeepEyesV2—across real-world understanding, OCR, chart interpretation, and mathematical reasoning tasks under three conditions: with tools, without tools, and using pure textual reasoning. Through ablation studies analyzing tool invocation traces, execution outcomes, and output formatting constraints, the study reveals that tool use does not consistently improve performance: 93%–96% of problems solved with tools could also be resolved without them, and tool invocation fails to substantially reduce computational or generative overhead. These findings highlight a significant disconnect between tool calling and genuine capability gains.

benchmark gainscapability evaluationmultimodal agents

Hot Scholars

HS

Harshit Surana

Co-Founder at Chaos Genius
Scientific DiscoveryMachine LearningConvex Optimization
MD

Mingyu Ding

Assistant Professor, UNC Chapel Hill
RoboticsEmbodied AIComputer Vision
JG

Jinyang Guo

The University of Sydney
Deep LearningEfficient MethodsEdge Computing
LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery