Score
Design and implement procedures and algorithms that partition work into batches—for training, inference, simulation, or batched numerical solves—covering minibatch construction, sampling strategies, batch layout, and memory-efficient organization for batched linear solves and simulations. Analyze and optimize batch size, composition, scheduling, memory layout, and data-access patterns to trade off throughput, latency, memory use, and objective-specific metrics (e.g., convergence speed, retrieval accuracy, solver throughput).
This work addresses the challenge of efficiently solving numerous related linear programming (LP) subproblems arising in mixed-integer programming (MIP), particularly in contexts such as strong branching and bound tightening, where traditional methods fail to exploit GPU parallelism effectively. The authors propose a GPU-oriented batched first-order optimization method, reformulating the primal–dual hybrid gradient algorithm into matrix–matrix operations to significantly enhance parallel efficiency. This study represents the first systematic application of batched first-order methods to LP subproblems within MIP solvers, advocating that GPUs should perform core computational tasks rather than merely assist CPU-based heuristics. The approach promotes deeper co-design between MIP algorithms and GPU architectures. Experimental results demonstrate that the proposed method outperforms conventional simplex solvers under specific problem scales and hardware configurations.
Existing optimizer benchmarks are often confined to a single batch size, overlooking the reliability of hyperparameter scaling rules across varying batch sizes. This study systematically investigates the dependence of optimizer performance on batch size through large-scale language model pretraining experiments, comparative evaluations of multiple optimizers, and extensive hyperparameter tuning. It provides the first empirical evidence that no universal Muon scaling rule generalizes across training configurations, and that the optimal optimizer shifts dynamically with batch size. By exposing the failure of prevailing scaling assumptions, this work demonstrates that optimizers must be independently selected and tuned for specific batch sizes to effectively enhance training efficiency.
This paper investigates the online scheduling problem on identical parallel machines with initial setup times: job processing times are unknown and revealed only upon completion; setup times are known monotonic functions of batched job sets; and batching is dynamically configurable. For four settings—single/multiple machines with/without preemption—we propose online algorithms based on dynamic batch partitioning, monotonicity modeling, and competitive analysis. All algorithms achieve an asymptotically optimal competitive ratio of $Theta(log n + log m)$, where $n$ is the number of jobs and $m$ is the number of machines. This result establishes, for the first time in this uncertain scheduling model, a unified theoretical performance bound that is provably optimal up to constant factors. It significantly advances the state-of-the-art in online batch scheduling theory by closing the gap between known upper and lower bounds under general monotonic setup costs and dynamic batching.
This paper theoretically investigates the stability and generalization of minibatch SGD and local SGD, addressing a key gap in existing work—its overemphasis on optimization error while neglecting formal generalization guarantees. We propose a novel “training-error-driven stability analysis” framework, introducing for the first time an expectation-variance decomposition into stability modeling to explicitly characterize how training error influences generalization. Our theoretical analysis establishes that, under overparameterization, both algorithms achieve optimal generalization risk bounds; moreover, their generalization error decreases linearly with the number of parallel workers—a property we term *linear generalization speedup*. This result transcends conventional analyses focused solely on optimization convergence rates. To our knowledge, this is the first generalization theory for large-scale distributed learning that simultaneously provides rigorous stability guarantees and provable linear speedup in generalization performance.
Existing batch Bayesian optimization (BO) methods suffer significant performance degradation as batch size increases, failing to fully exploit parallel computing resources. To address this scalability bottleneck, we propose a novel paradigm for large-scale parallel BO: it decomposes the high-dimensional search space into orthogonal, axis-aligned low-dimensional subspaces and introduces the first-of-its-kind Expected Subspace Improvement (ESI) acquisition function, which jointly optimizes diverse yet convergent batch query points within each subspace. This approach effectively balances exploration and exploitation while enabling scalable parallelization. Empirical evaluation on standard benchmarks demonstrates that our method substantially outperforms sequential BO in wall-clock time and consistently achieves state-of-the-art or competitive performance among seven leading batch BO algorithms. The implementation is publicly available in MATLAB, confirming both efficiency and practical applicability.
本文研究了在串行批处理机上最小化加权延迟工作总量的问题,提出了伪多项式时间动态规划算法,并针对特殊情况开发了专门算法。
研究通过对比实验验证了minibatch持久性在节省数据上的有效性,特别是在大批次大小下,但对速度和能耗无明显改善。
This work addresses the limitations of conventional Gaussian processes in Bayesian optimization, which struggle to model unobservable inter-batch variations, exhibit poor generalization, and suffer from low data efficiency. To overcome these challenges, the study introduces System-Aware Neural ODE Processes (SANODEP) into the Bayesian optimization framework for the first time, integrating meta-learning to construct a prior model capable of effectively capturing time-varying stochastic batch dynamics. The proposed approach substantially enhances both generalization performance and sample efficiency under both in-distribution and out-of-distribution batch conditions. Demonstrated on a penicillin batch production case study, the method achieves superior optimization outcomes with only a small number of experimental trials, yielding better objective values more rapidly than traditional Gaussian process-based approaches and significantly accelerating the initial optimization phase.
This work addresses the solution of general linear systems via the Randomized Block Kaczmarz (RBK) method, focusing on convergence analysis and performance enhancement. We propose a unified non-expansive block Kaczmarz framework that, for the first time, establishes rigorous convergence theory for diverse static random sampling strategies. By leveraging concentration inequalities and scale-invariant analysis techniques, we derive tight, expectation-based linear convergence rate bounds—substantially improving upon existing results. These bounds more accurately characterize the practical convergence behavior of block methods while maintaining computational efficiency even with low per-iteration cost. Extensive numerical experiments validate both the tightness of the theoretical bounds and the practical efficacy of the proposed algorithm.
This work addresses the lack of a unified and comparable benchmark for fairly evaluating rule-based, learning-based, and large language model (LLM)-driven autoscaling strategies in big data batch processing scenarios. To this end, we propose BatchBench, an open-source, workload-aware benchmarking framework. BatchBench introduces a taxonomy encompassing six representative batch workload types, features a parameterized workload generator whose fidelity is validated via two-sample Kolmogorov–Smirnov tests and Earth Mover’s Distance, and defines a five-dimensional evaluation protocol covering cost, SLA compliance, responsiveness, scaling jitter, and interpretability. Notably, it enables, for the first time, side-by-side comparison of all three autoscaling strategy categories while incorporating LLM inference cost accounting. The framework’s design is complete, and its reference implementation will be open-sourced to establish a standardized experimental foundation for autoscaling research.