Score
Designs and implements systems that partition and orchestrate a model’s runtime context and activations across multiple devices or processes to enable parallel, scalable inference and context-parallel execution. This work includes strategies for splitting activations and state, synchronizing and scheduling transfers, preserving pretrained weights during distributed inference, and providing context management/orchestration primitives and APIs.
A systematic analysis of large language model (LLM) inference system architectures—and the underlying technical synergies among their components—remains lacking. Method: We propose the first unified analytical framework that uncovers three foundational principles: workload forecasting, adaptive scheduling, and cost-aware compression. We introduce a deployment-paradigm-based taxonomy—categorizing systems into single-replica, multi-replica, decoupled, and serverless configurations—and comprehensively integrate key techniques including CUDA kernel optimization, continuous batching, PagedAttention, KV cache compression and persistence, weight/activation quantization, and memory offloading. Contribution/Results: Our work establishes the first holistic architecture map of LLM inference systems, explicitly characterizing inter-technique synergies and fundamental trade-offs. The resulting framework provides a systematic design guide for deploying LLMs efficiently, elastically, and cost-effectively.
Existing distributed large language model (LLM) inference in resource-constrained, highly heterogeneous, and dynamically evolving edge computing environments—such as Multi-access Edge Computing (MEC)—faces multiple challenges: volatile workloads, fluctuating bandwidth, node-level congestion, and real-time evolution of privacy constraints. To address these, this paper proposes the first adaptive sharding architecture tailored for LLM inference at the edge. It enables runtime capacity-aware node selection, operational-condition-driven dynamic model partition redistribution, and layer-granular real-time repartitioning. Leveraging dynamic workload modeling and a joint QoS-and-privacy-aware orchestration mechanism, the framework achieves co-optimization of low latency, high throughput, and strong privacy guarantees. Evaluated on realistic MEC deployments, our approach improves inference throughput by 42.3% and resource utilization by 35.7%, while rigorously satisfying heterogeneous QoS and privacy SLA requirements.
To address resource dynamics, multi-objective constraints (latency, utilization, privacy), and infrastructure instability in large language model (LLM) inference within heterogeneous edge environments, this paper proposes a runtime-reconfigurable framework for joint model partitioning and deployment optimization. It is the first to formulate layer-granular model partitioning and device placement as a dynamic constrained optimization problem, integrating model-aware capacity analysis, dynamic graph neural network–based repartitioning, and resource forecasting. Evaluated in a 6G multi-access edge computing scenario, the approach reduces end-to-end latency by 27.4% and improves average GPU utilization by 39.1% over static baselines, while enabling privacy-sensitive layers to execute locally. The core contribution lies in an online, constraint-adaptive inference scheduler that ensures theoretical rigor and practical deployability under time-varying operational conditions.
本文探讨了从本地优化到分布式控制的大型语言模型推理问题,通过结合vLLM和llm-d层,提出了一种推理执行计划器来优化执行策略。
This work addresses the orchestration bottlenecks faced by ultra-large-scale Sim-AI workflows on leadership-class supercomputers, which arise from task heterogeneity and extreme ensemble sizes. To overcome these challenges, the authors propose EnsembleLauncher, a recursively hierarchical and fully decentralized workflow orchestrator that introduces a decentralized control plane and a programmable scheduling policy interface, thereby surpassing conventional tools in both scalability and scheduling flexibility. Experiments on the Aurora supercomputer demonstrate that EnsembleLauncher can efficiently schedule system-wide resources to support up to 8 million serial tasks, achieving more than a fourfold performance improvement over state-of-the-art alternatives. Furthermore, it significantly enhances resource utilization for workloads with high task variance and active learning pipelines.
This work addresses the high latency in multi-agent systems caused by multi-step reasoning and redundant agent invocations during parallel execution, which often fails to meet real-time requirements. To this end, the authors propose LAMaS, a novel framework that introduces explicit latency supervision signals into multi-agent orchestration for the first time. LAMaS employs a learning-driven controller to construct an execution topology graph and leverages critical path analysis to optimize parallel scheduling. This approach departs from conventional paradigms centered on task performance or cost, instead prioritizing latency reduction along the critical path. Experimental results demonstrate that LAMaS reduces critical path length by 38%–46% compared to state-of-the-art methods across multiple benchmarks, while maintaining or even improving task performance.
Existing multimodal inference systems face challenges including rigid workflow orchestration, inefficient intermediate data transfer, and difficulty in sharing KV caches and model weights across heterogeneous components. This work proposes the first system-level unified abstraction for multimodal inference, featuring a three-tier architecture that co-optimizes control flow, data flow, and compute flow. The control flow layer employs a Python DSL to support both static and dynamic workflow orchestration; the data flow layer implements a zero-copy, distributed paged KV cache spanning GPUs, CPUs, and SSDs; and the compute flow layer enables multimodal prefix matching and KV reuse, unifying the forward passes of LLMs and diffusion models through a common SGLang interface. This design decouples orchestration logic from data transmission mechanisms, achieving efficient inference and resource reuse across diverse scenarios such as LongCat-Next dialogue and HunyuanImage-3 generation.
This work addresses the trade-off between efficiency and performance in multi-agent large language model inference, where existing approaches lack a unified framework for modeling diverse parallelization strategies. The study systematically distinguishes and unifies two inference-time parallelism mechanisms: replica parallelism—enabling multi-path exploration at the task level—and structural parallelism—supporting concurrent execution within a single reasoning path. To harmonize these strategies under a coherent execution semantics, the authors propose TIPEX, a novel coordination framework. Experimental results on the GAIA benchmark demonstrate that TIPEX significantly improves accuracy while reducing latency, with the most pronounced gains observed on medium-difficulty tasks. Notably, the findings also reveal that excessive parallelism can be detrimental, underscoring the necessity of tailoring parallelization strategies to the specific characteristics of each task.
This study addresses the limitation of existing test-time scaling methods, wherein models cannot autonomously determine context allocation and reuse. To overcome this, we propose Hermes, a framework that transfers context management decisions from fixed architectures to the model itself. Through a two-stage training paradigm, Hermes enables the model to acquire adaptive reasoning strategies, thereby achieving dynamically optimized allocation for multi-window computation. Experimental results demonstrate that our approach significantly enhances the performance of smaller models while exhibiting strong generalization capabilities across diverse benchmarks. Furthermore, Hermes reveals promising scaling potential that extends beyond its training compute budget, suggesting that empowering models with autonomous context management offers an effective pathway for test-time scaling.
该研究针对移动异构推理任务,提出一种分区感知调度方法,通过在线迭代搜索框架优化DAG调度,实现低延迟和高效执行。
This work addresses the challenges faced by large language model (LLM) agents operating over flat tool registries—namely, combinatorial explosion in decision space, context saturation, and degraded routing accuracy. To overcome these limitations, the authors propose a skill-tree-based hierarchical architecture that separates routing logic at internal nodes from execution at leaf nodes. Inspired by pushdown automata, the framework incorporates a LIFO stack-frame memory model and a lazy capability discovery mechanism, enabling isolated execution paths and scalable context management. The approach supports manifest-driven single-step execution loops and formal state modeling, significantly improving routing accuracy while reducing memory footprint and prompt costs under conditions of tool proliferation, multi-step workflows, and prompt exposure. This design meets enterprise-grade requirements for isolation and scalability.