Score
Designs, implements, and analyzes measurement systems and experiments that quantify model training cost and duration — e.g., per-batch and per-epoch wall‑clock time and compute consumed — to identify bottlenecks and hotspots, compare training time across models and settings, and produce standardized compute/time metrics and reports.
Rapid growth in distributed DNN model size far outpaces hardware evolution, making it increasingly difficult to design training systems that simultaneously achieve high efficiency and sustainability. Method: This paper introduces the first three-dimensional evaluation framework—spanning workload abstraction, simulation infrastructure, and total cost of ownership (TCO)/carbon emission modeling—grounded in systematic literature review, multi-dimensional comparative modeling, and quantitative analysis. Contribution/Results: We identify common limitations of existing simulators in workload characterization, resource modeling, and environmental impact assessment; propose a structured capability comparison matrix to clarify assumptions, functional boundaries, and applicability of each tool; and expose cross-layer modeling gaps, distilling key open research challenges. Our framework provides a reproducible, extensible benchmark and decision-support foundation for co-designing efficient, low-carbon distributed training systems.
This study systematically evaluates the impact of model quantization on the correctness and resource efficiency of deep learning systems, while also exploring methodologies for cross-study evidence aggregation in data-driven empirical research. Methodologically, it innovatively applies Structured Synthesis Methods (SSM) for the first time in this domain, integrating findings from six empirical studies covering 19 models through a qualitative-quantitative mixed analysis. Results demonstrate that quantization yields substantial resource gains—average storage compression of ×3.2, inference latency reduction of −41%, and GPU energy consumption decrease of −38%—with only a marginal correctness degradation (−1.7% on average), representing a well-controlled trade-off. The study identifies both consistent patterns and fragmentation bottlenecks in quantization effects, and proposes a refined empirical research framework and methodological guidelines tailored to quantization techniques. These contributions provide foundational methodological support and practical guidance for optimizing trustworthy AI systems.
In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.
This study addresses the challenge of quantifying and comparing large language model (LLM) behaviors across vendors in a standardized, cost-effective manner. We propose a simple, inexpensive, and reproducible framework for investigating model behavior by applying a fixed set of public stimuli across a cross-vendor panel of models. The framework innovatively integrates three complementary evaluation methods—exact matching, LLM-judge codebooks, and instrumented environments—to enable scalable behavioral tracking at minimal cost. Experiments reveal lexical convergence among models, evolving robustness to suffix-based prompts, divergences in stance adherence, and patterns of documentation non-compliance exhibited by coding agents. Collectively, this work establishes a systematic evaluation paradigm for tracking the behavioral evolution of large language models.
Massive, dynamic data streams in digital platforms render conventional ML monitoring methods ineffective or prohibitively costly in manual effort, forcing enterprises to downgrade to simpler models. Method: This paper proposes the Machine Learning Monitoring Agent (MLMA) framework, introducing a test-driven, automated retraining mechanism based on data-adaptive reference loss batches—designed to enable efficient closed-loop operations while preserving human-in-the-loop collaborative governance. The approach integrates design science principles, dynamic reference loss computation, key metric visualization, and human–AI collaborative workflows. Contribution/Results: Evaluated on a large-scale instant-delivery platform, MLMA supports concurrent monitoring of hundreds of models, significantly reduces manual intervention frequency, and sustains long-term online model performance stability. Its core contribution lies in unifying dynamic data adaptation, automated trigger logic, and human–AI collaboration—thereby overcoming critical technical bottlenecks in real-time monitoring and adaptive maintenance of large-scale ML systems.
Existing deep learning training energy estimation methods rely on unvalidated assumptions, resulting in high estimation errors. Method: This paper empirically uncovers the co-dependent energy consumption patterns between model architectures and hardware environments through multi-dimensional temporal monitoring (power, FLOPs, compute throughput, accuracy, etc.), regression modeling, energy-efficiency trade-off analysis, and cross-platform experiments. Contribution/Results: We propose four interpretable, high-accuracy energy estimation algorithms—reducing average error by 50%. For the first time, we establish a quantitative relationship among architecture, hardware environment, energy consumption, and model correctness. We discover that GPU selection should dynamically match model computational complexity to maximize energy efficiency; joint optimization achieves 80.72% energy reduction with <0.1% accuracy degradation. Our findings invalidate the reliability assumptions underpinning mainstream estimation approaches, and we open-source a lightweight energy prediction toolchain.
研究通过访谈13位机器学习从业者,分析了从笔记本原型到生产系统转换过程中涉及的工程变更及软件质量挑战,提出了监督债务的概念。
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.
This work addresses the challenge of systematically evaluating concept bottleneck models, whose applicability and failure mechanisms remain poorly understood due to the scarcity of real-world datasets with annotated concept labels. To bridge this gap, we introduce the first controllable synthetic benchmark that leverages parametric generation techniques to precisely modulate data modality, concept selection, annotation quality, and label completeness, thereby simulating diverse real-world relationships between concepts and predictions. This benchmark enables comprehensive evaluation of various concept bottleneck models across both decision-support and fully automated tasks, effectively identifying key performance determinants and characteristic failure modes. Our framework fills a critical void in the current evaluation landscape for concept-based interpretability methods.
研究通过审计Praxa AI管道文件及记录,发现代理评估指标与标签含义不一致的问题,并提供了一个可重用的验证包来区分不同类型的性能声明。