trajectory candidate generation

Designs and implements methods that synthesize multiple candidate step-by-step solution or action trajectories—expressed as chains-of-thought or multimodal sequences—using language models (including multimodal LMs). Builds tooling to produce, score, and select among those candidate trajectories by evaluating plausibility, goal alignment, and diversity.

trajectorycandidategeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.64
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users

Aug 24, 2025
SS
Sadia Sultana Chowa
🏛️ Daffodil International University | United International University | Charles Darwin University

Existing research on large language models (LLMs) as autonomous agents and tool users remains fragmented and limited in architecture design, multi-agent coordination, tool integration, cognitive mechanism modeling, and evaluation frameworks. Method: This survey systematically analyzes 2023–2025 top-tier conference and journal publications using structured literature analysis, integrating prompt engineering and fine-tuning techniques to dissect LLM implementations of core cognitive capabilities—reasoning, planning, and memory. Contribution/Results: We identify three breakthrough directions—verifiable reasoning, self-improvement, and personalized customization—and distill ten concrete future research pathways. Further, we propose a unified evaluation framework covering 68 publicly available datasets, exposing critical gaps in current benchmarks regarding task generalization, dynamic adaptability, and causal attribution capability.

Analyzing architectural designs for single and multi-agent systemsEvaluating benchmarks and datasets for LLM agent performanceExamining LLMs as autonomous agents and tool users

Must-Read Papers

Most classic and influential ideas
View more

Large language models for behavioral modeling: A literature survey

Sep 29, 2025
ML
Muhammad Laiq
🏛️ Mid Sweden University

A systematic literature review on the application of large language models (LLMs) to behavioral modeling—particularly automated generation of use case and sequence diagrams—is currently lacking, hindering research consolidation and practical guidance. Method: This paper presents the first comprehensive survey in this domain, identifying 14 core studies via a terminology-driven search strategy and synthesizing prevalent LLM application patterns and evaluation methodologies for behavioral modeling. Results: Findings confirm the feasibility of LLMs for diagram generation tasks; however, existing work is heavily reliant on GPT-series models and largely omits validation by domain experts. The study innovatively advocates for cross-model comparative analysis and expert-in-the-loop evaluation. It thereby provides theoretical foundations and actionable pathways for future research, tool development, and pedagogical practice in model-driven engineering and AI-assisted software modeling.

Assessing sequence diagram creation using language modelsEvaluating automated generation of use case diagramsSurveying LLM applications in behavioral modeling research

Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.

AI integrationMultimodal Large Language Modelsmultimodal perception

This work addresses the limitation of current large language models in autonomous tool use, which stems from a scarcity of diverse and realistic multi-turn tool interaction data. The authors propose a novel paradigm that automatically synthesizes multi-turn tool-use trajectories from general-purpose text corpora, treating natural text as a scalable source of behavioral traces for the first time. Their approach employs a four-stage pipeline—comprising relevance filtering, workflow and tool extraction, trajectory embodiment, and complexity optimization—alongside a dedicated trajectory synthesis model fine-tuned with supervised learning to enable efficient and generalizable data generation. Evaluated on the BFCL V3 multi-turn benchmark, the resulting GEM-32B model achieves a 16.5% performance gain, surpassing certain models trained on domain-specific τ-bench data while significantly reducing inference latency and computational cost.

autonomous agentsdata synthesislarge language models

Existing LLM-based tool learning methods predominantly formulate multi-step tool invocation as a text generation task, relying on supervised fine-tuning and thus struggling with the dynamic decision-making complexity inherent in sequential tool use. This work proposes the first step-wise reinforcement learning framework that explicitly models tool calling as a serialized decision process. We introduce a step-level reward shaping mechanism that separately quantifies the success and task-relevant contribution of each individual tool call. Further, we integrate policy gradient optimization with LLM–tool interface alignment to enable fine-grained policy updates. Evaluated on multi-step tool-use benchmarks, our approach achieves substantial improvements: +18.7% in task completion rate and +22.3% in tool-call accuracy, while significantly enhancing cross-step decision robustness. This framework establishes a novel paradigm for advancing LLM-based embodied intelligence and complex, multi-stage task execution.

Addressing decision-making complexities in tool learningEnhancing multi-step tool usage in LLMsImproving tool-use capabilities with step-grained reinforcement learning

Demystifying Chains, Trees, and Graphs of Thoughts

Jan 25, 2024
MB
Maciej Besta
🏛️ ETH Zurich | Dell | Cledar | BASF SE

Existing structured prompting paradigms—such as Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT)—lack a unified theoretical foundation, suffering from conceptual conflation and an absence of systematic taxonomy. Method: We propose the first comprehensive taxonomy for structured prompting, formally defining the notion of “reasoning topology,” constructing its spatial representation, and unifying CoT, ToT, and GoT through pipeline-based execution analysis, structural modeling, behavioral interpretation, and cross-paradigm empirical comparison. Contribution/Results: (1) We establish the first principled taxonomy for structured-prompt reasoning; (2) we uncover intrinsic relationships between topological structure and both reasoning performance and computational cost; and (3) we provide a theoretically grounded framework and design principles for scalable, interpretable prompt engineering.

Analyze performance and cost of different prompting designsDevelop taxonomy for structure-enhanced LLM reasoning schemesEnhance LLM reasoning with structured prompt engineering

Latest Papers

What's happening recently
View more

MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning

Oct 09, 2025
TA
Tajamul Ashraf
🏛️ Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) | University of Oxford

To address the scarcity of high-quality multimodal agent trajectories and the prohibitive cost of manual annotation, this paper proposes a vision-centric fine-tuning framework for Vision-Language Models (VLMs) as autonomous agents. Our method introduces three key innovations: (1) construction of M-TRACE, a large-scale, diverse multimodal task dataset; (2) Pref-X, a novel automated pipeline for synthesizing fine-grained, scalable multimodal preference pairs; and (3) an end-to-end optimization strategy integrating trajectory synthesis, behavioral cloning, and stepwise preference learning to jointly refine the VLM controller. Evaluated on three challenging benchmarks—Agent-X, GTA, and GAIA—our approach achieves state-of-the-art performance, significantly outperforming both leading open-source and proprietary VLMs in tool-use accuracy and cross-task robustness.

Addresses scarcity of high-quality multimodal trajectories for tool-use reasoningAutomatically synthesizes multimodal trajectories and generates preference pairsTrains VLM controllers for robust tool-use reasoning across benchmarks

This work addresses the modality gap in existing agent systems, which struggle to effectively integrate textual knowledge with parametric skills, thereby limiting task performance. To bridge this divide, the paper proposes treating model weights as a novel modality natively amenable to reasoning by large language models (LLMs), unifying parametric skills and textual knowledge into a composable representation for the first time. The approach employs prefix tuning to construct parametric skills, designs an LLM-augmented architecture, and introduces an instruction-driven weight composition mechanism to enable cross-modal skill generation and transfer. Experimental results demonstrate that the method significantly outperforms baselines relying solely on text or weights, achieving performance gains in multitask settings that are unattainable by single-modality approaches.

LLM reasoningmodality integrationparametric skills

Model-First Reasoning LLM Agents: Reducing Hallucinations through Explicit Problem Modeling

Dec 16, 2025
GK
Gaurav Kumar
🏛️ Stanford AI Professional Program | IESE EMBA Program

Large language models (LLMs) frequently exhibit high constraint violation rates and solution inconsistency in multi-step planning tasks due to implicit state tracking. To address this, we propose Model-First Reasoning (MFR), a two-stage paradigm: first, explicitly modeling problem entities, states, actions, and constraints—thereby integrating structured representations from classical AI planning into LLM reasoning; second, generating constraint-aware plans grounded in this explicit model. This design reveals that hallucination primarily stems from representational incompleteness, not inherent reasoning deficits. Extensive experiments across five domains—including medical scheduling and path planning—demonstrate that MFR reduces average constraint violation rates by 42% over Chain-of-Thought and ReAct, while significantly improving solution quality. Ablation studies confirm that explicit modeling is the primary source of performance gain, substantially enhancing planning robustness and interpretability.

Addresses representational deficiencies in LLM planning failuresImproves solution quality through explicit problem modelingReduces constraint violations in multi-step planning tasks

This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.

Hyper-parameter TuningLarge Language ModelsModeling and Simulation

Hot Scholars

YB

Yikun Ban

Beihang University, University of Illinois Urbana-Champaign
Reinforcement LearningEnsemble Learning
SL

Shiyi Lan

NVIDIA
VisionLLM AgentVisual Gen
YZ

Yakun Zhu

Shanghai Jiao Tong University
HG

Hao Geng

Harvard University
Theoretical Physics