compute-aware evaluation

Designs and implements evaluation protocols and benchmarks that control for compute by matching model exposure to tokens or other compute units — for example, scheduling checkpoints at uniform token intervals, enforcing token budgets, or aligning runs by cumulative compute. Builds metrics and analysis pipelines that prioritize interval-level performance (learning curves, per-token or per-interval efficiency) over endpoint-only comparisons to assess training and development efficiency under fixed compute budgets.

compute-awareevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.9
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current evaluations of large language models commonly employ fixed computational budgets during inference, which inadequately capture their true capabilities on complex tasks. This work systematically investigates the impact of inference-time computational resources—including token budgets, context compression, and repeated submissions—on model performance across seven challenging benchmarks. Using a unified evaluation framework applied to multiple state-of-the-art models, the study reveals for the first time that the allocation of inference computation significantly influences assessment outcomes: larger token budgets consistently enhance performance across diverse domains, whereas fixed budgets systematically underestimate the capabilities of advanced models. The authors argue that model competence should be conceptualized as a function of inference-time computation and advocate for transparent reporting of evaluation protocols.

benchmarkingfrontier modelsinference compute

This work addresses a critical flaw in existing LLM inference benchmarks, where single-process clients under high concurrency suffer from Python’s Global Interpreter Lock (GIL), leading to severe distortion in Time-to-First-Token (TTFT) and Time Per Output Token (TPOT) metrics. The study is the first to model the client as an M/G/1 queue, uncovering systematic bias introduced by queuing effects. To rectify this, the authors propose a multi-process, unbiased evaluation framework accompanied by a normalized metric—Normalized Time Per Output Token (NTPOT). This approach effectively eliminates client-side bottlenecks, enabling accurate and reproducible performance evaluation at scale, supporting thousands of queries per second. It substantially reduces the latency overestimation—often several-fold—in conventional benchmarks, thereby reflecting the true performance of production-grade LLM services.

benchmarkingLLM inferencemeasurement bias

Towards an Optimized Benchmarking Platform for CI/CD Pipelines

Oct 21, 2025
NJ
Nils Japke
🏛️ Technische Universität Berlin | DATEV eG

Performance regression detection in large-scale software systems is hindered by the high overhead and low frequency of traditional benchmarking, limiting its integration into CI/CD pipelines. This paper introduces CloudBench, an efficient performance benchmarking platform designed for cloud-native CI/CD. Its core contributions are: (1) composable lightweight optimizations—including sampling, differential execution, and cache reuse—that drastically reduce benchmarking overhead; (2) an automated regression detection mechanism combining statistical hypothesis testing with SLA-aware thresholds; and (3) a highly available, declarative architecture enabling seamless integration with mainstream CI/CD toolchains. Experimental evaluation demonstrates that CloudBench achieves 99% detection accuracy while reducing average benchmark execution time by 7.3×, thereby enabling per-commit performance validation. To our knowledge, CloudBench provides the first production-ready, systematic solution for continuous performance engineering.

Detecting performance regressions in CI/CD pipelines earlyIntegrating benchmark optimizations into practical CI/CD systemsOptimizing resource-intensive benchmarks for efficient execution

When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications

Apr 16, 2025
SH
Soren Henning
🏛️ Dynatrace Research | Johannes Kepler University Linz

This study addresses the high variability and low reliability of performance benchmarking results for stream-processing applications in cloud environments. Over three months, we conducted a large-scale longitudinal empirical study across multiple geographic regions and heterogeneous hardware—including diverse CPU architectures. Leveraging Kubernetes-based automated deployment, high-frequency repeated benchmarking, and time-series statistical analysis, we systematically characterized end-to-end cloud performance variability for the first time at the application level. We discovered that variability exhibits statistically significant diurnal and weekly periodicity (amplitude ≤2.5%) and a coefficient of variation <3.7%—substantially lower than commonly assumed in industry. Moreover, infrastructure sharing incurs at most a 2.5-percentage-point loss in measurement precision. These findings demonstrate strong robustness across regions and CPU architectures, providing empirical evidence and methodological foundations for enhancing reproducibility and trustworthiness in cloud-native benchmarking.

Assess temporal effects on stream processing applicationsEvaluate benchmark result accuracy across cloud regionsQuantify cloud performance variability impact on benchmarks

Latest Papers

What's happening recently
View more

This study addresses the challenges of service coordination and data auditing in message-passing experiments on high-performance computing (HPC) clusters by proposing a Slurm-orchestrated benchmarking framework. Methodologically, it introduces an "eligibility-first" evaluation paradigm that enforces data validity verification prior to throughput assessment, ensuring results satisfy causal logic constraints. Technically, the framework integrates a single-broker Kafka architecture, in-memory log storage, and multi-stage repeated validation mechanisms to achieve recoverable experiment control and distributed auditing. Experimental evaluations across 120 workloads yield 99 eligible observations, with selected workloads achieving a 100% qualification rate. Furthermore, the system attains a balanced endpoint throughput of 2,927 MiB/s with a P99 latency of 2.56 seconds, demonstrating both robust auditability and high performance.

Experiment QualificationHigh-Performance ComputingKafka Evaluation

This work addresses the synchronization bottleneck in Expert Parallel (EP) Mixture-of-Experts (MoE) inference, where layer-wise synchronization is constrained by the slowest GPU. Existing load-balancing strategies fail because they overlook the nonlinear variation of expert execution times across memory- and compute-intensive regimes. The paper is the first to characterize this bimodal behavior and proposes a makespan-aware scheduling approach. It models expert execution using a max-affine time model coupled with a phase-diagram prediction mechanism, formulates per-batch scheduling as a fixed-cost makespan minimization problem, and designs an efficient solver for real-time adaptive scheduling. Experiments demonstrate up to 15.5% higher throughput under mixed workloads, 4–6% end-to-end throughput improvement on Qwen3-235B, and approximately 15.6% reduction in p99 latency, with the phase diagram accurately predicting deployment outcomes.

expert-parallelload balancingmakespan

This work addresses the emerging risk that state-of-the-art large language models (LLMs) may detect and circumvent external control interventions—such as trajectory modifications—thereby undermining AI safety mechanisms. The study introduces the first systematic definition and quantification of “control intervention awareness” (CI-awareness), along with CIAware-Bench, a benchmark spanning four domains: argumentative writing, BigCodeBench, Bash Arena, and SHADE-Arena. Leveraging trajectory watermarking, auxiliary tasks, and diverse control protocols, the authors evaluate CI-awareness across 11 leading models via binary classification accuracy. Results reveal generally low to moderate CI-awareness (peak accuracy 0.87 versus random baseline 0.5), with higher detectability across model families. These findings indicate that CI-awareness is not an intrinsic property but rather contingent on task domain, model lineage, and deployment context, highlighting the presence of vendor-specific behavioral signatures.

AI controlcontrol interventionintervention awareness

This study addresses the challenges of fragmented academic computing resources and cross-facility large model pre-training by proposing a distributed training framework that integrates multiple global supercomputing centers. Methodologically, building upon the DiLoCo dual-loop architecture, it introduces elastic Nesterov outer-step optimization, a DARL heartbeat-based data leasing protocol, and a privilege-free queue-aware placement mechanism to enable resilient aggregation and efficient scheduling of intercontinental resources. Experimental results demonstrate that the system achieves fault-tolerant pre-training with zero data loss while reducing overhead to 3.1%, significantly shortening training cycles. This work provides an efficient and viable new paradigm for decentralized, large-scale LLM training.

Cross-facility pre-trainingDistributed LLM trainingFragmented compute allocations

Hot Scholars

HH

Huaibo Huang

NLPR, MAIS, CASIA
Computer VisionGenerative ModelsLow-level VisionFace Recognition
HH

Heng Huang

Brendan Iribe Endowed Professor in Computer Science, University Maryland College Park
Machine LearningAIBiomedical Data ScienceComputer Vision
DD

Devleena Das

Member of Technical Staff, AMD
Machine LearningOptimizationExplainability
GX

Guoxuan Xia

Imperial College London
efficient deep learninguncertainty estimation