Score
Designs and implements methods that synthesize multiple candidate step-by-step solution or action trajectories—expressed as chains-of-thought or multimodal sequences—using language models (including multimodal LMs). Builds tooling to produce, score, and select among those candidate trajectories by evaluating plausibility, goal alignment, and diversity.
This work investigates the fundamental applicability boundaries of large language models (LLMs) in automated planning. Through a systematic literature review, a multi-dimensional capability assessment framework, and empirical evaluation on canonical domains—including Block World and Logistics—the study reveals critical limitations: inconsistent long-horizon reasoning, failure in constraint-sensitive planning, and unreliable state tracking. Methodologically, it employs rigorous comparative analysis across diverse planning tasks to isolate intrinsic LLM deficiencies. The primary contribution is the first principled argument that LLMs are unsuitable as standalone planners; instead, it proposes “hybrid intelligent planning”—a novel paradigm wherein LLMs serve exclusively as semantic understanding and heuristic generation modules, tightly integrated with symbolic reasoning engines and search algorithms. The work establishes a reproducible, taxonomy-based evaluation methodology and provides concrete architectural design principles for synergistic LLM–symbolic system integration, thereby delivering both theoretical foundations and practical guidelines for LLM-augmented planning.
Existing research on large language models (LLMs) as autonomous agents and tool users remains fragmented and limited in architecture design, multi-agent coordination, tool integration, cognitive mechanism modeling, and evaluation frameworks. Method: This survey systematically analyzes 2023–2025 top-tier conference and journal publications using structured literature analysis, integrating prompt engineering and fine-tuning techniques to dissect LLM implementations of core cognitive capabilities—reasoning, planning, and memory. Contribution/Results: We identify three breakthrough directions—verifiable reasoning, self-improvement, and personalized customization—and distill ten concrete future research pathways. Further, we propose a unified evaluation framework covering 68 publicly available datasets, exposing critical gaps in current benchmarks regarding task generalization, dynamic adaptability, and causal attribution capability.
A systematic literature review on the application of large language models (LLMs) to behavioral modeling—particularly automated generation of use case and sequence diagrams—is currently lacking, hindering research consolidation and practical guidance. Method: This paper presents the first comprehensive survey in this domain, identifying 14 core studies via a terminology-driven search strategy and synthesizing prevalent LLM application patterns and evaluation methodologies for behavioral modeling. Results: Findings confirm the feasibility of LLMs for diagram generation tasks; however, existing work is heavily reliant on GPT-series models and largely omits validation by domain experts. The study innovatively advocates for cross-model comparative analysis and expert-in-the-loop evaluation. It thereby provides theoretical foundations and actionable pathways for future research, tool development, and pedagogical practice in model-driven engineering and AI-assisted software modeling.
Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.
This work addresses the limitation of current large language models in autonomous tool use, which stems from a scarcity of diverse and realistic multi-turn tool interaction data. The authors propose a novel paradigm that automatically synthesizes multi-turn tool-use trajectories from general-purpose text corpora, treating natural text as a scalable source of behavioral traces for the first time. Their approach employs a four-stage pipeline—comprising relevance filtering, workflow and tool extraction, trajectory embodiment, and complexity optimization—alongside a dedicated trajectory synthesis model fine-tuned with supervised learning to enable efficient and generalizable data generation. Evaluated on the BFCL V3 multi-turn benchmark, the resulting GEM-32B model achieves a 16.5% performance gain, surpassing certain models trained on domain-specific τ-bench data while significantly reducing inference latency and computational cost.
Existing LLM-based tool learning methods predominantly formulate multi-step tool invocation as a text generation task, relying on supervised fine-tuning and thus struggling with the dynamic decision-making complexity inherent in sequential tool use. This work proposes the first step-wise reinforcement learning framework that explicitly models tool calling as a serialized decision process. We introduce a step-level reward shaping mechanism that separately quantifies the success and task-relevant contribution of each individual tool call. Further, we integrate policy gradient optimization with LLM–tool interface alignment to enable fine-grained policy updates. Evaluated on multi-step tool-use benchmarks, our approach achieves substantial improvements: +18.7% in task completion rate and +22.3% in tool-call accuracy, while significantly enhancing cross-step decision robustness. This framework establishes a novel paradigm for advancing LLM-based embodied intelligence and complex, multi-stage task execution.
Existing structured prompting paradigms—such as Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT)—lack a unified theoretical foundation, suffering from conceptual conflation and an absence of systematic taxonomy. Method: We propose the first comprehensive taxonomy for structured prompting, formally defining the notion of “reasoning topology,” constructing its spatial representation, and unifying CoT, ToT, and GoT through pipeline-based execution analysis, structural modeling, behavioral interpretation, and cross-paradigm empirical comparison. Contribution/Results: (1) We establish the first principled taxonomy for structured-prompt reasoning; (2) we uncover intrinsic relationships between topological structure and both reasoning performance and computational cost; and (3) we provide a theoretically grounded framework and design principles for scalable, interpretable prompt engineering.
To address the scarcity of high-quality multimodal agent trajectories and the prohibitive cost of manual annotation, this paper proposes a vision-centric fine-tuning framework for Vision-Language Models (VLMs) as autonomous agents. Our method introduces three key innovations: (1) construction of M-TRACE, a large-scale, diverse multimodal task dataset; (2) Pref-X, a novel automated pipeline for synthesizing fine-grained, scalable multimodal preference pairs; and (3) an end-to-end optimization strategy integrating trajectory synthesis, behavioral cloning, and stepwise preference learning to jointly refine the VLM controller. Evaluated on three challenging benchmarks—Agent-X, GTA, and GAIA—our approach achieves state-of-the-art performance, significantly outperforming both leading open-source and proprietary VLMs in tool-use accuracy and cross-task robustness.
This work addresses the modality gap in existing agent systems, which struggle to effectively integrate textual knowledge with parametric skills, thereby limiting task performance. To bridge this divide, the paper proposes treating model weights as a novel modality natively amenable to reasoning by large language models (LLMs), unifying parametric skills and textual knowledge into a composable representation for the first time. The approach employs prefix tuning to construct parametric skills, designs an LLM-augmented architecture, and introduces an instruction-driven weight composition mechanism to enable cross-modal skill generation and transfer. Experimental results demonstrate that the method significantly outperforms baselines relying solely on text or weights, achieving performance gains in multitask settings that are unattainable by single-modality approaches.
Large language models (LLMs) frequently exhibit high constraint violation rates and solution inconsistency in multi-step planning tasks due to implicit state tracking. To address this, we propose Model-First Reasoning (MFR), a two-stage paradigm: first, explicitly modeling problem entities, states, actions, and constraints—thereby integrating structured representations from classical AI planning into LLM reasoning; second, generating constraint-aware plans grounded in this explicit model. This design reveals that hallucination primarily stems from representational incompleteness, not inherent reasoning deficits. Extensive experiments across five domains—including medical scheduling and path planning—demonstrate that MFR reduces average constraint violation rates by 42% over Chain-of-Thought and ReAct, while significantly improving solution quality. Ablation studies confirm that explicit modeling is the primary source of performance gain, substantially enhancing planning robustness and interpretability.
This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.