test-time scaling orchestration

Designs and implements runtime orchestration components that choose and apply scaling strategies for multi-agent and orchestration-level systems (e.g., number of sub-agents, rollout depth, token budgets) to meet accuracy, latency, and resource-cost constraints. Builds and runs scalability experiments and evaluation frameworks, and engineers policies that select scaling decisions using orchestration rewards or heuristics to optimize task performance while avoiding costly sub-agent rollouts at runtime.

test-timescalingorchestration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.84
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work addresses the lack of a unified framework in current LLM agent workflows, which hinders method comparison and reproducibility. To resolve this, we propose the Agent Computation Graph (ACG) framework, which models workflows as computation graphs and adopts “structure determines timing” as a core principle. The framework explicitly distinguishes between reusable templates, runtime instance graphs, and execution traces, enabling a systematic categorization of static and dynamic optimization approaches. Through a comprehensive literature review and conceptual modeling, we develop a multidimensional evaluation framework that integrates structural properties, establishes precise terminology, and defines standardized evaluation criteria. This foundation supports a reproducible and highly comparable research paradigm for optimizing LLM agent workflows.

agentic computation graphsLLM agentsstatic vs dynamic workflows

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges faced by large language model (LLM) agents operating over flat tool registries—namely, combinatorial explosion in decision space, context saturation, and degraded routing accuracy. To overcome these limitations, the authors propose a skill-tree-based hierarchical architecture that separates routing logic at internal nodes from execution at leaf nodes. Inspired by pushdown automata, the framework incorporates a LIFO stack-frame memory model and a lazy capability discovery mechanism, enabling isolated execution paths and scalable context management. The approach supports manifest-driven single-step execution loops and formal state modeling, significantly improving routing accuracy while reducing memory footprint and prompt costs under conditions of tool proliferation, multi-step workflows, and prompt exposure. This design meets enterprise-grade requirements for isolation and scalability.

context window saturationdecision-space explosionLLM agents

Existing evaluation methods struggle to disentangle the quality of task orchestration in multi-agent systems from confounding factors such as agent capabilities and environmental noise, while real-world execution incurs prohibitive costs. To address this, this work proposes OrchBench—a deterministic simulation-based benchmarking platform that models task dependencies via directed acyclic graphs and enables isolated, efficient assessment of orchestration plans. OrchBench achieves the first interpretable evaluation of orchestration quality with dramatically reduced overhead: requiring only 1.3% of the tokens and 10.3% of the time compared to real execution, while maintaining high fidelity (Pearson r = 0.816). Furthermore, it reveals that information retention rate is more critical to performance than simply increasing the number of agents.

coordination overheaddeterministic simulationevaluation benchmark

This work addresses the current lack of open-source infrastructure capable of efficiently training and evaluating large-scale agents on complex tasks such as software engineering and computer operation. To this end, we propose a three-service decoupled architecture tailored for agent-environment interaction workloads, which separates the system into three independent services—model, agent, and environment—enabling fine-grained task scheduling, dynamic resource allocation, and unified interface communication. This design allows each component to scale independently and configure resources flexibly, significantly improving training efficiency and resource utilization. Experimental results demonstrate that the system can stably support tens of thousands of concurrent agent tasks, thereby filling a critical gap in infrastructure for large-scale agent training.

agent-environment interactionagentic AIdistributed orchestration

Existing approaches to automatic multi-agent system design suffer from complex orchestration, limited global reasoning capabilities, and ill-defined boundaries of their advantages. This work proposes MAS-Orchestra, a novel framework that formulates multi-agent orchestration as a function-call-based global reinforcement learning problem, enabling end-to-end generation of complete systems in a single pass. To systematically evaluate task characteristics, the authors introduce MASBENCH, a benchmark assessing performance across five dimensions: Depth, Horizon, Breadth, Parallelism, and Robustness. Experimental results demonstrate consistent performance gains on tasks such as mathematical reasoning and multi-hop question answering. Furthermore, the study reveals that the benefits of multi-agent collaboration are not universal but critically depend on task structure, verification mechanisms, and individual agent capabilities.

agent orchestrationefficacy evaluationmulti-agent systems

This work addresses the high latency in multi-agent systems caused by multi-step reasoning and redundant agent invocations during parallel execution, which often fails to meet real-time requirements. To this end, the authors propose LAMaS, a novel framework that introduces explicit latency supervision signals into multi-agent orchestration for the first time. LAMaS employs a learning-driven controller to construct an execution topology graph and leverages critical path analysis to optimize parallel scheduling. This approach departs from conventional paradigms centered on task performance or cost, instead prioritizing latency reduction along the critical path. Experimental results demonstrate that LAMaS reduces critical path length by 38%–46% compared to state-of-the-art methods across multiple benchmarks, while maintaining or even improving task performance.

inference latencylatencymulti-agent systems

Latest Papers

What's happening recently
View more

This study addresses the challenge of ensuring traceability, controllability, and correctness of large language model–driven agents within business processes while preserving their autonomy. To this end, the work proposes the first multidimensional attribute classification framework specifically designed for agent orchestration, integrating principles from business process management to strike a balance between agent autonomy and system robustness. Complementing the framework, the authors introduce qualitative decision-making guidelines and quantitative evaluation metrics. The efficacy of the proposed approach is empirically validated through multi-agent experiments in a predictive light-sensing scenario, demonstrating its capacity to support both theoretical inquiry and practical deployment of orchestrated intelligent agents in real-world applications.

Agentic OrchestrationAutonomyBusiness Process Management

This work proposes the first large language model (LLM)-driven agent framework for autonomous, end-to-end management of high-performance computing (HPC) applications in cloud environments. Addressing the heavy reliance on manual intervention and the lack of intelligent decision-making in traditional HPC cloud deployment, the framework enables automated multi-platform container construction, Kubernetes-based orchestration, cross-instance performance optimization, and adaptive elastic scaling policy generation. By integrating LLM-powered agents into HPC cloud workflow orchestration, this study establishes a novel paradigm of automation and self-adaptation. Experimental evaluation across four representative HPC applications demonstrates that the system achieves expert-level linear scalability, substantially reduces job completion time, and yields actionable best practices for collaborative agent design in HPC contexts.

Agentic OrchestrationCloud ComputingHPC Applications

This work addresses the orchestration bottlenecks faced by ultra-large-scale Sim-AI workflows on leadership-class supercomputers, which arise from task heterogeneity and extreme ensemble sizes. To overcome these challenges, the authors propose EnsembleLauncher, a recursively hierarchical and fully decentralized workflow orchestrator that introduces a decentralized control plane and a programmable scheduling policy interface, thereby surpassing conventional tools in both scalability and scheduling flexibility. Experiments on the Aurora supercomputer demonstrate that EnsembleLauncher can efficiently schedule system-wide resources to support up to 8 million serial tasks, achieving more than a fourfold performance improvement over state-of-the-art alternatives. Furthermore, it significantly enhances resource utilization for workloads with high task variance and active learning pipelines.

exascaleorchestration bottlenecksscalability

This study addresses the challenge of automating workflows in complex industries—such as logistics, healthcare, and construction—where processes are fragmented across heterogeneous tools and involve multi-party collaboration. The work proposes orchestration as a core abstraction to enable effective automation by dynamically coordinating multi-step tasks, enforcing domain-specific constraints, managing human approvals, and integrating legacy systems. It introduces the novel concept of “orchestration bottlenecks” and develops a theoretical framework that unifies multi-agent systems, workflow modeling, constraint reasoning, and human–AI collaboration, while exposing critical gaps in current multi-agent approaches at the orchestration level. Based on distinct sources of operational friction across domains, the paper advocates for targeted architectural safeguards—such as constraint enforcement or explainability—and phased implementation strategies to provide actionable pathways for automation in complex operational environments.

legacy systemsoperationally complex industriesorchestration

Reinforcement learning (RL) has long remained absent from real-world service orchestration deployments, commonly attributed to telemetry latency, load fluctuations, and tenant-induced unpredictability—yet systematic empirical validation is lacking. This study rigorously evaluates performance degradation of three representative RL-based orchestration systems under production-grade perturbations, employing preregistered experiments, paired statistical inference, and family-wise error correction. Results reveal that most performance gains claimed in the literature either fail to replicate or diminish substantially; only one system consistently outperforms Kubernetes Horizontal Pod Autoscaler (HPA) by approximately 40× under observational delay. The findings indicate that existing advantages often stem from inadequate baselines, limited artifacts, or evaluation biases, underscoring an urgent need for deployment-oriented evaluation standards and institutional incentives to bridge the gap between research and practice.

evaluation biasinstitutional incentivesproduction deployment

Hot Scholars

TH

Torsten Hoefler

Professor of Computer Science at ETH Zurich
High Performance ComputingDeep LearningNetworkingMessage Passing Interface
JH

Jia-Huei Ju

University of Amsterdam
Natural Language ProcessingInformation Retrieval
AY

Andrew Yates

Johns Hopkins University, Human Language Technology Center of Excellence
Information RetrievalNLPAI
JX

Jiawei Xue

Purdue University; Alibaba Group
LLMsGNNsRecommendationUrban Science