Score
Designs and implements traversal and clustering workflows that use a large language model to make semantic judgments and guide navigation over items or nodes, including systems that discover clusters via LLM-provided similarity or grouping signals. Builds traversal orders and decision logic (for example, minimum spanning tree–based traversals) that minimize and selectively trigger LLM calls, invoking the model only for complex, uncertain, or high-value cases to reduce cost and latency.
Large language model (LLM) systems commonly suffer from substantial resource waste due to static or suboptimal deployment strategies. Method: This paper proposes a cost-quality–aware query routing mechanism that dynamically dispatches user queries to the most suitable lightweight model, domain-specific expert, or embedding strategy. We formally define the routing problem for the first time and introduce a novel taxonomy jointly optimizing relevance and resource efficiency. Our framework systematically compares academic approaches with industrial practices, incorporating query understanding, policy selection, multi-granularity routing (at both model and embedding levels), fine-grained cost modeling, and a unified evaluation protocol. Contribution/Results: Experiments demonstrate that our mechanism significantly reduces inference overhead—by up to 42% in latency and 38% in GPU memory—while maintaining or even improving answer quality across diverse benchmarks. This work establishes both a theoretical foundation and a reproducible, practical paradigm for building efficient, scalable, and cost-effective LLM systems.
Current LLM-based algorithm design relies heavily on empirical trial-and-error, lacking formal theoretical foundations to systematically analyze how critical design choices—such as task decomposition strategies and prompt engineering—affect accuracy and computational efficiency. Method: We propose the first formal analytical framework for LLM-invocation algorithms, modeling LLM subroutines as a computational graph, establishing structured principles for task decomposition, and introducing an error propagation model that enables provable analysis of both accuracy and computational complexity. Contribution/Results: Empirically validated across parallel, hierarchical, and recursive paradigms, our framework explains observed empirical phenomena, guides prompt design and granularity selection, predicts performance bottlenecks, and inspires novel robust algorithm designs. The implementation is publicly available.
Static deployment of large language models (LLMs) struggles to dynamically select the optimal model based on query complexity and domain, leading to an imbalance between performance and cost. This work proposes a three-dimensional framework—centered on decision timing, information sources, and computation mechanisms—to systematically analyze dynamic routing and cascading strategies across multiple LLMs. It introduces the first taxonomy of routing paradigms spanning independently trained LLMs and integrates diverse technical approaches, including query difficulty estimation, human preference modeling, uncertainty quantification, reinforcement learning, and multimodal fusion. Experimental results demonstrate that well-designed routing systems can surpass the strongest individual model, achieving a superior trade-off between inference efficiency and performance, while also highlighting critical challenges such as generalization.
研究使用大型语言模型设计针对库存控制、排队网络控制和商品优化等问题的近似最优算法,通过两种不同层次的应用展示了模型的有效性。
Large language models (LLMs) struggle with implicit planning–oriented agent tasks due to heavy reliance on extensive tool integration, manual prompt engineering, or costly fine-tuning. Method: This paper proposes explicitly modeling domain-specific procedural knowledge as Hierarchical Task Networks (HTNs), integrating both handcrafted and LLM-generated HTNs into the reasoning process to guide task decomposition and execution. Contribution/Results: Experiments demonstrate that HTN augmentation significantly improves task success rates—20B/70B LLMs outperform a 120B baseline, and handcrafted HTNs enable smaller models to surpass larger ones, confirming that knowledge-driven structuring can meaningfully offset architectural scale disadvantages. This work formally establishes HTNs as an effective mechanism for enhancing LLM-based agents, revealing the critical role of structured procedural knowledge in agent design. It introduces a new paradigm for building lightweight, interpretable, and high-performance LLM agents grounded in explicit task hierarchies.
This study addresses the lack of systematic understanding regarding the practical usage patterns, reliability mechanisms, and autonomy levels of large language model (LLM) agents in low-code/no-code platforms. Drawing on over 6,000 publicly available n8n workflows, the authors employ large-scale data mining, structured log analysis, and qualitative coding to empirically characterize how LLM agents are deployed in real-world automation scenarios—specifically examining task distribution, workflow structure, tool invocation, and degrees of autonomy. The findings reveal that while LLMs are commonly embedded within complex workflows featuring control logic and human review steps, such workflows generally lack structured fault tolerance, repair loops, and approval mechanisms. Based on these insights, the study articulates ten empirical observations and five design implications to inform the development of more reliable and governable low-code platforms.
本文通过动态构建和优化任务导向的本体来解决大语言模型在特定领域任务中的知识利用不全和多步决策脆弱的问题。
This work proposes a data-driven, end-to-end approach to automatically construct and optimize large language model (LLM) workflows, addressing the deployment bottlenecks associated with manual pipeline design. The workflow construction is formulated as a bilevel optimization problem: the outer loop searches over high-level structural configurations, while the inner loop performs differentiable optimization of individual LLM invocation modules using textual gradients, enabling layer-wise adjustments analogous to backpropagation. This is the first method to integrate bilevel optimization with textual gradients, allowing efficient workflows to be discovered fully automatically without human intervention. Experimental results demonstrate that the proposed approach achieves performance on par with or superior to strong baseline systems that rely on either handcrafted or automatically generated workflows across multiple tasks.
This work addresses the limitation of large language models (LLMs) in reasoning tasks, where despite having access to the full search history, their implicit trajectory representation hinders effective backtracking and reuse of past states. To overcome this, the paper introduces LinTree, a novel method that explicitly encodes a linearized search tree structure within the LLM’s reasoning trajectory, using parent pointers to clearly delineate backtracking paths. This enables the model to efficiently identify and reuse previously explored search states. Experimental results on benchmark tasks—including Blocks World, grid navigation, and Sokoban—demonstrate that LinTree significantly outperforms both implicit reasoning and heuristic-based LLM search approaches, achieving higher task success rates and improved search efficiency. These findings underscore the effectiveness of explicit structural representations in enhancing LLM reasoning capabilities.
"This study addresses the challenge of transforming natural language routing rules, authored by business administrators, into executable workflow diagrams for enterprise contact centers. The project employs a neural-symbolic decomposition approach, utilizing a compact intermediate representation and a deterministic compiler to minimize the need for direct graph construction by large language models. A key innovation is the integration of a learned registry selection front-end, which enhances the generation of relevant vocabulary, significantly improving the quality and efficiency of Directed Acyclic Graph (DAG) workflow generation. Implemented across four distinct models, the system achieved an effectiveness score of approximately 89%, with condition accuracy around 90% and JSON format correctness rates of 99%-100%. Additionally, the token count per rule was reduced by half compared to monolithic prompting methods."