Score
Designs and implements evaluation protocols and benchmarks that control for compute by matching model exposure to tokens or other compute units — for example, scheduling checkpoints at uniform token intervals, enforcing token budgets, or aligning runs by cumulative compute. Builds metrics and analysis pipelines that prioritize interval-level performance (learning curves, per-token or per-interval efficiency) over endpoint-only comparisons to assess training and development efficiency under fixed compute budgets.
该研究通过ContinuumBench基准解决了云-边缘环境中服务放置与自动扩展联合评估的问题,采用控制实验方法比较了不同控制器在多种条件下的表现。
Current evaluations of large language models commonly employ fixed computational budgets during inference, which inadequately capture their true capabilities on complex tasks. This work systematically investigates the impact of inference-time computational resources—including token budgets, context compression, and repeated submissions—on model performance across seven challenging benchmarks. Using a unified evaluation framework applied to multiple state-of-the-art models, the study reveals for the first time that the allocation of inference computation significantly influences assessment outcomes: larger token budgets consistently enhance performance across diverse domains, whereas fixed budgets systematically underestimate the capabilities of advanced models. The authors argue that model competence should be conceptualized as a function of inference-time computation and advocate for transparent reporting of evaluation protocols.
This work addresses a critical flaw in existing LLM inference benchmarks, where single-process clients under high concurrency suffer from Python’s Global Interpreter Lock (GIL), leading to severe distortion in Time-to-First-Token (TTFT) and Time Per Output Token (TPOT) metrics. The study is the first to model the client as an M/G/1 queue, uncovering systematic bias introduced by queuing effects. To rectify this, the authors propose a multi-process, unbiased evaluation framework accompanied by a normalized metric—Normalized Time Per Output Token (NTPOT). This approach effectively eliminates client-side bottlenecks, enabling accurate and reproducible performance evaluation at scale, supporting thousands of queries per second. It substantially reduces the latency overestimation—often several-fold—in conventional benchmarks, thereby reflecting the true performance of production-grade LLM services.
Performance regression detection in large-scale software systems is hindered by the high overhead and low frequency of traditional benchmarking, limiting its integration into CI/CD pipelines. This paper introduces CloudBench, an efficient performance benchmarking platform designed for cloud-native CI/CD. Its core contributions are: (1) composable lightweight optimizations—including sampling, differential execution, and cache reuse—that drastically reduce benchmarking overhead; (2) an automated regression detection mechanism combining statistical hypothesis testing with SLA-aware thresholds; and (3) a highly available, declarative architecture enabling seamless integration with mainstream CI/CD toolchains. Experimental evaluation demonstrates that CloudBench achieves 99% detection accuracy while reducing average benchmark execution time by 7.3×, thereby enabling per-commit performance validation. To our knowledge, CloudBench provides the first production-ready, systematic solution for continuous performance engineering.
This study addresses the high variability and low reliability of performance benchmarking results for stream-processing applications in cloud environments. Over three months, we conducted a large-scale longitudinal empirical study across multiple geographic regions and heterogeneous hardware—including diverse CPU architectures. Leveraging Kubernetes-based automated deployment, high-frequency repeated benchmarking, and time-series statistical analysis, we systematically characterized end-to-end cloud performance variability for the first time at the application level. We discovered that variability exhibits statistically significant diurnal and weekly periodicity (amplitude ≤2.5%) and a coefficient of variation <3.7%—substantially lower than commonly assumed in industry. Moreover, infrastructure sharing incurs at most a 2.5-percentage-point loss in measurement precision. These findings demonstrate strong robustness across regions and CPU architectures, providing empirical evidence and methodological foundations for enhancing reproducibility and trustworthiness in cloud-native benchmarking.
This study addresses the challenges of service coordination and data auditing in message-passing experiments on high-performance computing (HPC) clusters by proposing a Slurm-orchestrated benchmarking framework. Methodologically, it introduces an "eligibility-first" evaluation paradigm that enforces data validity verification prior to throughput assessment, ensuring results satisfy causal logic constraints. Technically, the framework integrates a single-broker Kafka architecture, in-memory log storage, and multi-stage repeated validation mechanisms to achieve recoverable experiment control and distributed auditing. Experimental evaluations across 120 workloads yield 99 eligible observations, with selected workloads achieving a 100% qualification rate. Furthermore, the system attains a balanced endpoint throughput of 2,927 MiB/s with a P99 latency of 2.56 seconds, demonstrating both robust auditability and high performance.
This work addresses the synchronization bottleneck in Expert Parallel (EP) Mixture-of-Experts (MoE) inference, where layer-wise synchronization is constrained by the slowest GPU. Existing load-balancing strategies fail because they overlook the nonlinear variation of expert execution times across memory- and compute-intensive regimes. The paper is the first to characterize this bimodal behavior and proposes a makespan-aware scheduling approach. It models expert execution using a max-affine time model coupled with a phase-diagram prediction mechanism, formulates per-batch scheduling as a fixed-cost makespan minimization problem, and designs an efficient solver for real-time adaptive scheduling. Experiments demonstrate up to 15.5% higher throughput under mixed workloads, 4–6% end-to-end throughput improvement on Qwen3-235B, and approximately 15.6% reduction in p99 latency, with the phase diagram accurately predicting deployment outcomes.
This work addresses the emerging risk that state-of-the-art large language models (LLMs) may detect and circumvent external control interventions—such as trajectory modifications—thereby undermining AI safety mechanisms. The study introduces the first systematic definition and quantification of “control intervention awareness” (CI-awareness), along with CIAware-Bench, a benchmark spanning four domains: argumentative writing, BigCodeBench, Bash Arena, and SHADE-Arena. Leveraging trajectory watermarking, auxiliary tasks, and diverse control protocols, the authors evaluate CI-awareness across 11 leading models via binary classification accuracy. Results reveal generally low to moderate CI-awareness (peak accuracy 0.87 versus random baseline 0.5), with higher detectability across model families. These findings indicate that CI-awareness is not an intrinsic property but rather contingent on task domain, model lineage, and deployment context, highlighting the presence of vendor-specific behavioral signatures.
研究通过改变计算条件测试AI模型在负责任的AI基准测试中的结论稳定性,使用不同批次、量化和基准缩减方法评估模型的准确性、偏见等性能。
This study addresses the challenges of fragmented academic computing resources and cross-facility large model pre-training by proposing a distributed training framework that integrates multiple global supercomputing centers. Methodologically, building upon the DiLoCo dual-loop architecture, it introduces elastic Nesterov outer-step optimization, a DARL heartbeat-based data leasing protocol, and a privilege-free queue-aware placement mechanism to enable resilient aggregation and efficient scheduling of intercontinental resources. Experimental results demonstrate that the system achieves fault-tolerant pre-training with zero data loss while reducing overhead to 3.1%, significantly shortening training cycles. This work provides an efficient and viable new paradigm for decentralized, large-scale LLM training.