probabilistic latency modeling

Design and validate probabilistic models that characterize the distribution of task or operation execution times and end-to-end latencies, explicitly accounting for stalls, contention, and interference from co-located workloads. Use those models to estimate deadline-miss probabilities and latency tails and to produce uncertainty-aware inputs for scheduling and resource-management decisions.

probabilisticlatencymodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Exploiting the Uncertainty of the Longest Paths: Response Time Analysis for Probabilistic DAG Tasks

Apr 02, 2025
YG
Yiyang Gao
🏛️ Sun Yat-sen University | University of York | Southeast University | City University of Hong Kong

Response-time analysis of p-DAG tasks in parallel real-time systems (e.g., autonomous driving) suffers from excessive conservatism and poor scalability—existing approaches are either overly pessimistic or rely on infeasible exhaustive enumeration. Method: This paper proposes an efficient analytical framework based on the probability distribution of the longest path. It is the first to jointly identify longest-path structures and compute their occurrence probabilities, rigorously modeling interference loads and scheduling semantics to derive a formally correct response-time distribution. Contribution/Results: The method eliminates scenario enumeration, reducing computational overhead by six orders of magnitude. On benchmark workloads, it achieves a mean absolute error of only 1.04%; over 90% of cases exhibit errors below 5%. This significantly improves timing guarantee accuracy and resource utilization efficiency while ensuring formal correctness.

Analyzing probabilistic response times in parallel real-time systemsOvercoming scalability issues in timing analysis for p-DAGsReducing computation cost while maintaining accuracy guarantees

How long can you sleep? Idle Time System Inefficiencies and Opportunities

Oct 08, 2025
GA
Georgia Antoniou
🏛️ University of Cyprus | Rivos Inc.

To address low utilization of deep idle states in latency-sensitive applications, this paper identifies a significant gap between theoretically available idle opportunities and their actual exploitation—caused by inaccurate idle scheduling decisions and non-negligible deep-sleep transition latency. We propose a queueing-theoretic modeling framework that integrates M/M/1, c×M/M/1, and M/M/c models, calibrated with real-world server workload traces, to quantify system-level idle potential under diverse configurations. For the first time, we systematically identify numerous untriggered deep-idle entry opportunities and develop a scalable methodology for idle-efficiency evaluation. The framework provides quantifiable, early-stage guidance for hardware–OS co-design, enabling energy-efficiency optimization and supporting system-level power management strategies that explicitly balance latency constraints and energy savings.

Identifying inefficiencies in entering deep idle power statesModeling idle time distribution in latency-critical server systemsProviding early-stage design exploration for server configurations

This study addresses a critical limitation in classical queueing analysis—its frequent neglect of preemption overhead—which hinders accurate assessment of stability and response time in preemptive scheduling systems. Focusing on the M/G/1 queue with preemption overhead, this work investigates class-based preemptive priority scheduling and presents the first exact analysis of response time distributions for such systems. By introducing a novel theoretical construct termed “task joint transform,” which integrates Laplace transforms with stochastic process techniques, the authors derive recursive formulas for the Laplace transforms of response times for tasks of arbitrary classes. This framework enables closed-form computation of all response time moments, clearly elucidates the performance impact of preemption overhead, and establishes a general analytical foundation extendable to broader scheduling overhead models.

M/G/1 queuepreemption overheadpriority scheduling

ML Inference Scheduling with Predictable Latency

Dec 21, 2025
HZ
Haidong Zhao
🏛️ Inria | Sorbonne University

ML inference serving on GPUs suffers from unpredictable latency due to resource interference under concurrent execution, making it challenging to simultaneously satisfy SLOs, meet deadline constraints, and achieve high GPU utilization. Existing interference prediction approaches are coarse-grained, static, and lack runtime adaptability. To address these limitations, this work proposes a fine-grained, dynamically adaptive interference prediction mechanism: it leverages measurement-driven co-location interference feature extraction, integrates online workload characterization, and employs a lightweight dynamic prediction model—all embedded within a closed-loop scheduling framework. Experimental evaluation demonstrates that our approach improves SLO compliance by 23.6%, reduces tail-latency variability by 41%, and sustains GPU utilization above 82% under multi-model colocation. Collectively, it significantly enhances scheduling predictability and resource efficiency.

Address unpredictable latency from GPU interference in ML inference schedulingDevelop adaptive models to handle diverse workload characteristics effectivelyImprove coarse-grained interference prediction by considering runtime co-location dynamics

Latest Papers

What's happening recently
View more

This work addresses the lack of resource-centric computational efficiency metrics—specifically in terms of node-hours—for existing supercomputers and large-scale AI training platforms operating under high failure rates. It proposes the first efficiency evaluation framework grounded in resource consumption rather than execution time, unifying failure rate, mean time between failures, and checkpoint/restart overhead into a cohesive resource-based model. The framework extends Daly’s (2006) model to accommodate heterogeneous scientific workloads. Validated on one year of production data from the Frontier supercomputer, the approach leverages runtime log analysis, joint modeling of failures and checkpointing, and optimization algorithms to accurately quantify the expected fraction of resources usable for scientific computation and to determine optimal checkpoint intervals that minimize resource loss.

application failurescomputational efficiencyExascale computing

Existing formal methods struggle to quantify the likelihood and latency of asynchronous, out-of-order orchestrations or effectively assess how resources—such as communication delays, computational overhead, and fault recovery—affect system behavior. This work proposes AsInst, a novel orchestration language that, for the first time, integrates probabilistic and temporal resource awareness into semantic modeling. It employs temporal Bayesian networks to capture runtime values and their availability times, and establishes semantic equivalence with Future-based network semantics. The approach enables joint analysis of execution likelihood and performance, supports Ozone-style selective merge constructs, and has been successfully applied to communication fault recovery and runtime performance prediction, demonstrating its effectiveness in evaluating latency and reliability.

asynchronous choreographiesout-of-order executionprobabilistic

This work addresses the high bias and variance commonly observed in A/B tests of datacenter scheduling policies, which arise due to Markovian interference. To tackle this challenge, the paper introduces a novel hybrid causal inference method that uniquely integrates Little’s Law with a Differences-in-Q estimator. This approach explicitly models key complexities inherent in real-world queueing systems, including non-stationary arrival rates, heterogeneous service rates, and communication delays. Theoretical analysis and extensive simulations demonstrate that the proposed method substantially reduces both estimation bias and variance, achieving superior robustness and higher accuracy across a variety of practical scenarios compared to existing approaches.

A/B testingLittle's LawMarkovian interference

Hot Scholars

JG

Jeremy Goldwasser

Statistics PhD Student, UC Berkeley
NLPInterpretabilityStatistical ML
GQ

Guanqiao Qu

The University of Hong Kong
Artificial IntelligenceMachine LearningEdge IntelligenceNetworking
XC

Xianhao Chen

Assistant Professor, The University of Hong Kong
Wireless networksmobile edge computingedge AIdistributed learning
QL

Qian Lin

Research Engineer, ByteDance
DatabaseDistributed SystemData Streams
WL

Weichen Liu

College of Computing and Data Science, Nanyang Technological University
Embedded SystemsMultiprocessor SystemsNetwork-on-Chip