Score
Designs and implements orchestration layers, APIs, and libraries that expose multiple algorithms and models behind a unified interface, enable algorithm- and model-agnostic inference and plug-and-play swapping of strategies without reconfiguring hardware, and manage routing, versioning, and execution of model inference. Builds integrations with external orchestration frameworks, provides unified programmatic (e.g., Python) and no-code human-in-the-loop interfaces, and analyzes system behavior to ensure correct orchestration, resource allocation, and interoperability across components.
This work addresses the challenges faced by large language model (LLM) agents operating over flat tool registries—namely, combinatorial explosion in decision space, context saturation, and degraded routing accuracy. To overcome these limitations, the authors propose a skill-tree-based hierarchical architecture that separates routing logic at internal nodes from execution at leaf nodes. Inspired by pushdown automata, the framework incorporates a LIFO stack-frame memory model and a lazy capability discovery mechanism, enabling isolated execution paths and scalable context management. The approach supports manifest-driven single-step execution loops and formal state modeling, significantly improving routing accuracy while reducing memory footprint and prompt costs under conditions of tool proliferation, multi-step workflows, and prompt exposure. This design meets enterprise-grade requirements for isolation and scalability.
This work addresses the fragmentation and incompatibility across multiple LLM providers in tool-calling interfaces, message formats, and streaming behaviors, which hinder the portability and reproducibility of agent systems. To resolve this, the authors propose Orchestral, a lightweight Python framework that enables cross-LLM agent development through a unified, type-safe interface abstraction. Its core innovations include automatic tool schema generation driven by Python type hints, a synchronous streaming execution model that balances determinism with interactivity, and a modular, decoupled architecture. Orchestral supports standardized message and tool representations, context compression, sandboxed workspaces, MCP integration, and sub-agent mechanisms, substantially reducing engineering complexity while enhancing system portability, maintainability, and functional completeness.
Current tool-augmented large language model (LLM) ecosystems suffer from fragmentation—characterized by coexisting heterogeneous protocols (e.g., OpenAI Function Calling, Toolformer), manual schema definition, and complex execution orchestration—leading to low development efficiency and high integration overhead. To address this, we propose a protocol-agnostic unified tool integration framework. Our approach introduces an abstract protocol layer for cross-standard compatibility, an automated schema inference mechanism to eliminate manual specification, and a dual-mode concurrent scheduler enabling seamless synchronous and asynchronous tool execution. Experimental evaluation demonstrates that, compared to baseline approaches, our framework reduces implementation code volume by 60–80%, achieves up to 3.1× improvement in end-to-end execution latency, and maintains full backward compatibility with mainstream LLM tool-calling ecosystems.
Existing multi-agent collaborative systems are hindered by static workflows, sequential scheduling, and heterogeneous interfaces, leading to high complexity and poor scalability. This work proposes Agent-as-Tool, a unified paradigm that abstracts both agents and tools into a standardized, learnable action space, and introduces ParaManager—a lightweight coordinator enabling state-aware parallel subtask decomposition, delegation, and asynchronous execution. By unifying communication protocols and incorporating explicit state feedback, the framework facilitates efficient multi-agent collaboration. A two-stage training strategy—combining supervised fine-tuning with a recovery mechanism and reinforcement learning—optimizes task success rate, protocol compliance, response diversity, and reasoning efficiency. Experiments demonstrate that ParaManager achieves strong performance across multiple benchmarks and exhibits robust generalization to unseen agent pools.
This work addresses the lack of structured, verifiable, and reusable decision mechanisms in existing automated machine learning approaches for model selection. It proposes a semantic task profiling–based structured agent framework that leverages retrieval-augmented generation of historical cases and code modules to construct an intermediate representation blueprint encompassing modeling components, composition logic, and execution constraints. By integrating code execution feedback with a failure-aware reinforcement learning strategy, the framework enables memory-driven, traceable, multi-stage search optimization. Evaluated on financial time-series forecasting and generation tasks, the method significantly outperforms both conventional AutoML systems and current agent-based baselines, achieving consistent improvements in task performance, execution success rate, and decision interpretability.
This study systematically investigates the capability boundaries of large language models (LLMs) in security tool orchestration, with a focus on the relative impact of model choice, client implementation, toolset composition, and reasoning mechanisms on system performance. Leveraging the open-source orchestration framework HexStrike-AI, the authors conduct multi-configuration comparative experiments across 86 picoCTF challenges, complemented by failure diagnosis and targeted refinements—including tool corrections, behavioral adjustments, and capability extensions—to quantitatively demonstrate, for the first time, the critical role of the client component in determining the performance of a fixed LLM. Results indicate that performance bottlenecks primarily stem from reasoning or environmental constraints rather than missing tools, enabling an increase in overall solve rate from 55.4% to 72.0% (p < 0.001) with high reproducibility (17 out of 20 trials consistent). The work introduces a reproducible evaluate-and-improve feedback loop, establishing a new paradigm for intelligent security agent systems.
This work addresses the lack of standardized human–AI collaborative interfaces in existing self-driving laboratories by introducing the Model Context Protocol (MCP) to this domain for the first time. The authors propose a unified software architecture based on an MCP server, through which all experimental functionalities are exposed via MCP, enabling both human users and AI agents to interact through a shared interface. The system supports automatic tool discovery and dynamically generates visual programming interfaces, significantly lowering the barrier to no-code experimental workflow design. The feasibility of this architecture is demonstrated in a color-matching self-driving laboratory case study, showing seamless compatibility with both manual human operation and AI agent invocation.
This work addresses the limitations of existing agent orchestration frameworks, which rely on external schedulers and incur substantial context overhead, require state-of-the-art large language models, and risk exposing proprietary workflows. To overcome these issues, the authors propose compiling multi-node agent workflows—comprising up to 55 nodes—directly into the weights of a small fine-tuned language model, thereby creating what they term “underground agents.” This approach provides the first systematic demonstration that complex workflows can be effectively internalized within model parameters. By integrating structured workflow representations, task-specific knowledge injection, and decision-hub modeling, the method achieves performance comparable to leading models on tasks such as travel booking, Zoom customer support, and insurance claims processing, while reducing inference costs by two orders of magnitude and substantially diminishing reliance on conventional orchestration frameworks.
This study addresses the challenge of automating workflows in complex industries—such as logistics, healthcare, and construction—where processes are fragmented across heterogeneous tools and involve multi-party collaboration. The work proposes orchestration as a core abstraction to enable effective automation by dynamically coordinating multi-step tasks, enforcing domain-specific constraints, managing human approvals, and integrating legacy systems. It introduces the novel concept of “orchestration bottlenecks” and develops a theoretical framework that unifies multi-agent systems, workflow modeling, constraint reasoning, and human–AI collaboration, while exposing critical gaps in current multi-agent approaches at the orchestration level. Based on distinct sources of operational friction across domains, the paper advocates for targeted architectural safeguards—such as constraint enforcement or explainability—and phased implementation strategies to provide actionable pathways for automation in complex operational environments.
This work addresses the orchestration bottlenecks faced by ultra-large-scale Sim-AI workflows on leadership-class supercomputers, which arise from task heterogeneity and extreme ensemble sizes. To overcome these challenges, the authors propose EnsembleLauncher, a recursively hierarchical and fully decentralized workflow orchestrator that introduces a decentralized control plane and a programmable scheduling policy interface, thereby surpassing conventional tools in both scalability and scheduling flexibility. Experiments on the Aurora supercomputer demonstrate that EnsembleLauncher can efficiently schedule system-wide resources to support up to 8 million serial tasks, achieving more than a fourfold performance improvement over state-of-the-art alternatives. Furthermore, it significantly enhances resource utilization for workloads with high task variance and active learning pipelines.