context parallelism

Designs and implements systems that partition and orchestrate a model’s runtime context and activations across multiple devices or processes to enable parallel, scalable inference and context-parallel execution. This work includes strategies for splitting activations and state, synchronizing and scheduling transfers, preserving pretrained weights during distributed inference, and providing context management/orchestration primitives and APIs.

contextparallelism

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Adaptive Orchestration for Inference of Large Foundation Models at the Edge

Mar 19, 2025
FK
Fernando Koch
🏛️ Florida Atlantic University | Technical University Munich | Carl Zeiss AG

Existing distributed large language model (LLM) inference in resource-constrained, highly heterogeneous, and dynamically evolving edge computing environments—such as Multi-access Edge Computing (MEC)—faces multiple challenges: volatile workloads, fluctuating bandwidth, node-level congestion, and real-time evolution of privacy constraints. To address these, this paper proposes the first adaptive sharding architecture tailored for LLM inference at the edge. It enables runtime capacity-aware node selection, operational-condition-driven dynamic model partition redistribution, and layer-granular real-time repartitioning. Leveraging dynamic workload modeling and a joint QoS-and-privacy-aware orchestration mechanism, the framework achieves co-optimization of low latency, high throughput, and strong privacy guarantees. Evaluated on realistic MEC deployments, our approach improves inference throughput by 42.3% and resource utilization by 35.7%, while rigorously satisfying heterogeneous QoS and privacy SLA requirements.

Adapting split inference to dynamic workloads and network conditionsBalancing latency, throughput, and privacy in edge AI applicationsOrchestrating Large Foundation Model inference in resource-constrained edge environments

Joint Partitioning and Placement of Foundation Models for Real-Time Edge AI

Nov 30, 2025
AD
Aladin Djuhera
🏛️ Technical University of Munich | Florida Atlantic University | Carl Zeiss AG

To address resource dynamics, multi-objective constraints (latency, utilization, privacy), and infrastructure instability in large language model (LLM) inference within heterogeneous edge environments, this paper proposes a runtime-reconfigurable framework for joint model partitioning and deployment optimization. It is the first to formulate layer-granular model partitioning and device placement as a dynamic constrained optimization problem, integrating model-aware capacity analysis, dynamic graph neural network–based repartitioning, and resource forecasting. Evaluated in a 6G multi-access edge computing scenario, the approach reduces end-to-end latency by 27.4% and improves average GPU utilization by 39.1% over static baselines, while enabling privacy-sensitive layers to execute locally. The core contribution lies in an online, constraint-adaptive inference scheduler that ensures theoretical rigor and practical deployability under time-varying operational conditions.

Addressing resource volatility in heterogeneous edge environmentsDynamically partitioning and placing foundation models for edge AI inferenceOptimizing real-time inference under latency, utilization, and privacy constraints

This work addresses the orchestration bottlenecks faced by ultra-large-scale Sim-AI workflows on leadership-class supercomputers, which arise from task heterogeneity and extreme ensemble sizes. To overcome these challenges, the authors propose EnsembleLauncher, a recursively hierarchical and fully decentralized workflow orchestrator that introduces a decentralized control plane and a programmable scheduling policy interface, thereby surpassing conventional tools in both scalability and scheduling flexibility. Experiments on the Aurora supercomputer demonstrate that EnsembleLauncher can efficiently schedule system-wide resources to support up to 8 million serial tasks, achieving more than a fourfold performance improvement over state-of-the-art alternatives. Furthermore, it significantly enhances resource utilization for workloads with high task variance and active learning pipelines.

exascaleorchestration bottlenecksscalability

This work addresses the high latency in multi-agent systems caused by multi-step reasoning and redundant agent invocations during parallel execution, which often fails to meet real-time requirements. To this end, the authors propose LAMaS, a novel framework that introduces explicit latency supervision signals into multi-agent orchestration for the first time. LAMaS employs a learning-driven controller to construct an execution topology graph and leverages critical path analysis to optimize parallel scheduling. This approach departs from conventional paradigms centered on task performance or cost, instead prioritizing latency reduction along the critical path. Experimental results demonstrate that LAMaS reduces critical path length by 38%–46% compared to state-of-the-art methods across multiple benchmarks, while maintaining or even improving task performance.

inference latencylatencymulti-agent systems

Latest Papers

What's happening recently
View more

Existing multimodal inference systems face challenges including rigid workflow orchestration, inefficient intermediate data transfer, and difficulty in sharing KV caches and model weights across heterogeneous components. This work proposes the first system-level unified abstraction for multimodal inference, featuring a three-tier architecture that co-optimizes control flow, data flow, and compute flow. The control flow layer employs a Python DSL to support both static and dynamic workflow orchestration; the data flow layer implements a zero-copy, distributed paged KV cache spanning GPUs, CPUs, and SSDs; and the compute flow layer enables multimodal prefix matching and KV reuse, unifying the forward passes of LLMs and diffusion models through a common SGLang interface. This design decouples orchestration logic from data transmission mechanisms, achieving efficient inference and resource reuse across diverse scenarios such as LongCat-Next dialogue and HunyuanImage-3 generation.

distributed KV cacheheterogeneous computingintermediate data transmission

This work addresses the trade-off between efficiency and performance in multi-agent large language model inference, where existing approaches lack a unified framework for modeling diverse parallelization strategies. The study systematically distinguishes and unifies two inference-time parallelism mechanisms: replica parallelism—enabling multi-path exploration at the task level—and structural parallelism—supporting concurrent execution within a single reasoning path. To harmonize these strategies under a coherent execution semantics, the authors propose TIPEX, a novel coordination framework. Experimental results on the GAIA benchmark demonstrate that TIPEX significantly improves accuracy while reducing latency, with the most pronounced gains observed on medium-difficulty tasks. Notably, the findings also reveal that excessive parallelism can be detrimental, underscoring the necessity of tailoring parallelization strategies to the specific characteristics of each task.

execution coordinationinference-time parallelismmulti-agent LLM systems

This study addresses the limitation of existing test-time scaling methods, wherein models cannot autonomously determine context allocation and reuse. To overcome this, we propose Hermes, a framework that transfers context management decisions from fixed architectures to the model itself. Through a two-stage training paradigm, Hermes enables the model to acquire adaptive reasoning strategies, thereby achieving dynamically optimized allocation for multi-window computation. Experimental results demonstrate that our approach significantly enhances the performance of smaller models while exhibiting strong generalization capabilities across diverse benchmarks. Furthermore, Hermes reveals promising scaling potential that extends beyond its training compute budget, suggesting that empowering models with autonomous context management offers an effective pathway for test-time scaling.

Context allocationContextual reasoningInference-time compute

This work addresses the challenges faced by large language model (LLM) agents operating over flat tool registries—namely, combinatorial explosion in decision space, context saturation, and degraded routing accuracy. To overcome these limitations, the authors propose a skill-tree-based hierarchical architecture that separates routing logic at internal nodes from execution at leaf nodes. Inspired by pushdown automata, the framework incorporates a LIFO stack-frame memory model and a lazy capability discovery mechanism, enabling isolated execution paths and scalable context management. The approach supports manifest-driven single-step execution loops and formal state modeling, significantly improving routing accuracy while reducing memory footprint and prompt costs under conditions of tool proliferation, multi-step workflows, and prompt exposure. This design meets enterprise-grade requirements for isolation and scalability.

context window saturationdecision-space explosionLLM agents

Hot Scholars

YW

Yongji Wu

UC Berkeley
Machine Learning SystemsDatacenter Networks
YS

Yogesh Simmhan

Associate Professor, Indian Institute of Science
Distributed SystemsEdge AcceleratorsGraph AnalyticsCloud Computing
HJ

Heng Ji

Professor of Computer Science, AICE Director, ASKS Director, UIUC, Amazon Scholar
Natural Language ProcessingLarge Language Models
XY

Xiaohu Yang

National University of Defense Technology
Plasma physicsLaser-plasma interactionInertial confinement fusionCharged particle beam
HC

Howard Chen

Princeton University
Machine LearningNatural Language Processing