Score
Designs, builds, and analyzes systems and pipelines that discover and detect signals in real time or batch, including writing and deploying detection rules, detection-as-code, and runtime detection systems; and develops statistical and ML-based detectors, automated detection pipelines, and model-driven detection components. Validates and measures detection performance—using detection and estimation methods and efficacy metrics—and operates detection rule engineering, discovery, and validation workflows for runtime and automated environments.
Massive, dynamic data streams in digital platforms render conventional ML monitoring methods ineffective or prohibitively costly in manual effort, forcing enterprises to downgrade to simpler models. Method: This paper proposes the Machine Learning Monitoring Agent (MLMA) framework, introducing a test-driven, automated retraining mechanism based on data-adaptive reference loss batches—designed to enable efficient closed-loop operations while preserving human-in-the-loop collaborative governance. The approach integrates design science principles, dynamic reference loss computation, key metric visualization, and human–AI collaborative workflows. Contribution/Results: Evaluated on a large-scale instant-delivery platform, MLMA supports concurrent monitoring of hundreds of models, significantly reduces manual intervention frequency, and sustains long-term online model performance stability. Its core contribution lies in unifying dynamic data adaptation, automated trigger logic, and human–AI collaboration—thereby overcoming critical technical bottlenecks in real-time monitoring and adaptive maintenance of large-scale ML systems.
Current IDEs lack intelligent, end-to-end support for the machine learning (ML) lifecycle, while MLOps platforms remain decoupled from coding environments. To bridge this gap, we propose a novel large language model (LLM)-enhanced intelligent IDE paradigm that deeply integrates LLMs into the development environment. This enables synergistic, closed-loop automation across code-level intelligent programming—such as code generation, debugging, and completion—and full-stack MLOps pipeline orchestration—including data validation, feature store management, data drift detection, retraining triggers, and CI/CD deployment. The system unifies development, experimentation, validation, and monitoring phases, significantly improving engineering efficiency and reproducibility. Empirical evaluation on the UCI Adult and M5 datasets demonstrates a 61% reduction in pipeline configuration time, a 45% improvement in experimental reproducibility, and a 14% increase in data drift detection accuracy.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This paper addresses the prevalent “silent failures” and unpredictable behavior of large language models (LLMs) when auto-generating code for embedded machine learning (ML) workflows. We propose a closed-loop evaluation framework covering data preprocessing, model conversion, and on-device inference code generation. Through multi-model empirical analysis, we introduce the first failure taxonomy for LLM-generated code in embedded ML, identifying systemic fragility arising from prompt format bias, implicit structural assumptions encoded in LLMs, and blind spots in compilation- and runtime-level validation. Key failure patterns include format-misleading parsing errors and “compilable-yet-functionally-broken” runtime errors—both largely undetectable by conventional verification methods. Our findings provide both theoretical foundations and practical guidelines for enhancing the reliability, traceability, and robustness of LLM-driven embedded ML systems.
Addressing the acute shortage of AI/ML expertise in software engineering (SE), this study investigates the effectiveness and adoption barriers of AutoML for SE decision-making. Method: We systematically benchmark 12 state-of-the-art AutoML tools (e.g., H2O, Auto-sklearn, TPOT) on SE datasets and complement quantitative evaluation with surveys and expert interviews. Contribution/Results: Our empirical analysis reveals that AutoML-generated models achieve significantly higher average accuracy than manually tuned models on SE classification tasks. However, 83% of the tools lack automated feature engineering and deployment capabilities, and provide insufficient workflow support for non-ML experts—exposing a critical “pseudo end-to-end” limitation. The study identifies structural gaps in full-lifecycle automation and cross-role collaboration within current AutoML systems, thereby providing evidence-based insights and concrete design directions for next-generation AutoML tailored to SE contexts.
研究通过SuriCap平台和CTF式工作坊,分析了60名参与者创建网络入侵检测规则的过程与方法,揭示经验对规则质量影响有限,并指出标记数据的重要性。
This study addresses sequential multi-stream detection under the constraint that only one data stream can be observed at each time step, with the goal of simultaneously controlling global false alarm and missed detection probabilities while minimizing detection delay. To this end, the work introduces a novel optimality criterion based on the expected order statistics of detection times and proposes an active sampling strategy—dubbed “follow-the-leader”—that integrates exploration and exploitation mechanisms. Theoretical analysis demonstrates that the proposed strategy achieves asymptotic optimality for all such criteria as error probabilities vanish. Numerical experiments further confirm its superior finite-sample performance compared to existing methods and show that it closely approaches the performance of an ideal oracle policy that has full knowledge of the anomalous streams.
研究探讨了AI辅助软件工程中,通过分层监督(包括预防性、可执行性和人工监督)来应对传统代码审查等控制机制面临的压力。
This study addresses the challenge of root-cause localization in automotive software testing, where high-dimensional sensor data generated during hardware-in-the-loop (HIL) simulations render traditional threshold-based methods ineffective. Existing data-driven approaches often require extensive labeled data and lack interpretability, failing to meet ISO 26262 traceability requirements. To overcome these limitations, the authors propose a two-stage diagnostic framework: first, safety requirements are automatically verified on a dSPACE real-time platform to filter anomalous test records; then, sliding windows of sensor signals are abstracted into statistical, relational, and contextual descriptors, which are fed as fixed prompts to an open-source large language model fine-tuned with 4-bit low-rank adaptation (LoRA). Evaluated on six fault types injected into a gasoline engine, the approach achieves 81.6% accuracy with a minimal 2B-parameter model—comparable to larger models—while operating entirely on a single consumer-grade GPU. This work pioneers the use of instruction-tuned large language models for sensor-level automotive fault diagnosis, demonstrating that diagnostic performance hinges more on task-specific adaptation convergence than on model scale, thereby achieving high accuracy, data efficiency, and explainable decision-making.
本文通过结合探索性分析与自动化分析方法,开发了一个实时监控BGP流量异常的系统,以帮助安全分析师更有效地检测网络安全威胁。