Score
Designs and implements systems and protocols for generating, assigning, initializing, and recording random seeds and independent parallel random streams, including explicit initialization, closed-form or analytic and asymptotic seed schedules per regime, and scalable assignment across workers. Builds and analyzes seed-controlled experiments and replicates—stratifying seeds, ensuring statistical independence, logging seed configurations for provenance, and measuring/reporting seed-induced variance and outcomes.
研究探讨了通过转移伪随机数生成器的完整内部状态来确保跨库一致性和可移植性的问题,使用Mersenne Twister和Philox在四个Python生态系统中进行实验,提出实现保真度是科学重现性的必要条件。
This work addresses the inefficiency of manual performance tuning in cloud-native stream processing systems, which heavily relies on expert experience. To automate and accelerate configuration optimization, the authors propose an experiment-driven approach that integrates Latin hypercube sampling, simulated annealing, and hill climbing into a three-stage search strategy. This method is deeply coupled with the Theodolite benchmarking framework to automatically orchestrate experiments on Kubernetes and preemptively terminate underperforming configurations. Evaluated on Kafka Streams, the approach efficiently explores the configuration space and identifies settings that substantially outperform default configurations, achieving up to a 23% improvement in throughput. The study demonstrates a practical and effective pathway toward automated, high-efficiency tuning of stream processing systems in cloud-native environments.
为解决间歇性和Lévy过程的分类与建模问题,IntLevPy提供Python库,通过参数估计、拟合优化、分类方法(如调整R²和Γ指标)及模拟验证实现。
Seed science faces interdisciplinary complexity, data scarcity, and a lack of standardized benchmarks—key barriers hindering the application of large language models (LLMs) in plant breeding. To address this, we propose SeedBench, the first multi-task evaluation benchmark specifically designed for seed science, covering core tasks including germplasm identification, hybrid prediction, genotype–phenotype inference, and breeding decision support. SeedBench integrates domain expertise to construct realistic task chains and a curated multimodal seed dataset, enabling zero-shot and few-shot evaluation. Comprehensive assessment across 26 state-of-the-art LLMs—including proprietary, open-source, and domain-adapted variants—reveals a substantial capability gap in agricultural reasoning for general-purpose models, while domain-tuned models demonstrate marked performance gains. SeedBench establishes a reproducible, extensible evaluation framework for agri-AI, providing both rigorous benchmarking standards and actionable pathways for advancing LLM-driven precision breeding.
This work addresses the risk of selecting suboptimal configurations under limited pretraining budgets, which can lead to significant resource waste. To mitigate this, the authors propose an auditable, staged promotion protocol that operates within a fixed micro-pretraining environment. By employing multi-stage time budgets ranging from 2 minutes to 12 hours, predefined promotion rules, replicated experiments across heterogeneous hardware (A100/L40S), and multiple random seeds, the method leverages Staged Factorial Screening and the val_bpb metric to distinguish genuine operational evidence of superiority from performance fluctuations due to single-seed variance. Strict near-equivalence and mean-gap criteria are applied to control cost and risk. The entire process consumes only 169.2 GPU-hours and successfully identifies a bridging configuration that consistently leads at the 12-hour stage, achieving over 60% GPU-hour savings compared to full-scale training.
研究提出一种决策层方法,通过统计意义的重用、生成或延迟策略,解决流系统中专家模型池的维护问题。
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
研究通过改变训练种子并固定数据分区,分析了推荐系统评估中随机性的影响,指出单次随机种子可能导致评估结论不稳定。
This study addresses the inherent bias in existing parallel control methods and their difficulty in achieving precise inference-time control by proposing the first exact parallel control algorithm. The approach employs a trajectory storage mechanism to circumvent irreversible process simulation and utilizes a truncation scheme to replace incomputable time reversal, thereby enabling generative model guidance without retraining while strictly preserving target distribution invariance. By integrating non-equilibrium replica exchange with stochastic Monte Carlo analysis, the method achieves both high precision and diversity across biomolecular sampling and image generation tasks. Furthermore, it demonstrates strong robustness when applied to distilled samplers.
"This study addresses the presence of bias in the unpatched probability samples within the MCP server ecosystem. It randomly selected 400 npm/stdio servers, probed using publicly available seeds, to evaluate their initialization handshake success rates, JSON Schema compliance, and tool description redundancy. For the first time, it reveals the true state of unpatched samples, identifying primary issues such as server startup failures and significant discrepancies in tool descriptions across benchmark sets. The methodology included random sampling, network probing, JSON Schema validation, and cosine similarity. The results indicate that only 48.8% of the servers successfully completed the initialization handshake; among the 195 operational servers, the tool-level omission rate was 58.8%; and the cross-author approximate duplication rate for genuine MCP tools was 0%, while for BFCL v4, it was 16.7%."