Score
Designs, implements, and iteratively refines computational models and their training pipelines, including choosing architectures or algorithms, preparing and transforming inputs (feature engineering), and implementing training, optimization, and hyperparameter tuning. Builds evaluation and validation workflows to measure performance, diagnose errors or bias, and ready models for deployment or further analysis.
Existing LLM-driven automated visual modeling approaches rely on global, one-shot optimization, resulting in poor attribution, slow convergence, low stability, and limited accessibility for non-experts. Method: We propose an “iterative single-component fine-tuning” strategy, inspired by expert human practice, wherein only one module in the pipeline is optimized per iteration. This is integrated with training-feedback-guided modular updates, zero-shot prompt engineering, and a multi-domain evaluation protocol to construct an end-to-end LLM agent framework. Contribution/Results: Our approach significantly enhances interpretability, stability, and convergence efficiency of optimization. Evaluated across multiple standard benchmarks and Kaggle datasets, it consistently outperforms state-of-the-art zero-shot LLM methods, achieving superior classification accuracy and generalization capability.
Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.
This work investigates whether model ensembles within the 1–3B parameter range can enhance code generation performance through execution feedback and pipeline architectures. We construct a generate-and-refine pipeline based on small language models, incorporate an execution feedback mechanism, and employ a NEAT-inspired evolutionary algorithm to search for optimal topologies. Our experiments reveal that execution feedback is pivotal—yielding performance gains exceeding four standard deviations on HumanEval and MBPP, primarily by correcting runtime errors—whereas increased topological complexity offers no significant benefit. The refinement component’s capability outweighs the identity of the generator, and single-run evaluations tend to overestimate evolutionary improvements; early stopping proves essential to prevent performance degradation. Moreover, specialized code models consistently outperform all combinations of general-purpose models.
Existing CAD generation models struggle to emulate engineers’ iterative design processes and lack the capability to validate physical and structural compliance. This work proposes an industry-native CAD generation framework that produces complete multi-part STEP files from engineering text and, for the first time, integrates finite element analysis (FEA) into the generative loop to verify structural plausibility. The approach leverages structured blueprint descriptions and 21-view image renderings as dual supervisory signals to guide large language model agents—such as GPT-5.5 and Claude Code—toward self-improving generation. Evaluated on the S2O and Fusion360 datasets, the method significantly enhances geometric reconstruction quality and engineering compliance, improving Box-IoU from 0.444 to 0.592 and from 0.397 to 0.505, respectively.
This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.
本文综述了AI驱动的科学计算工作流在编排、执行、可重复性和来源方面的问题,并提出了解决这些问题的方法和系统需求。
研究通过访谈13位机器学习从业者,分析了从笔记本原型到生产系统转换过程中涉及的工程变更及软件质量挑战,提出了监督债务的概念。
本文提出两种技术,通过增加推理时间和模拟真实部署环境来提高对齐评估的真实性,解决模型在测试与实际部署中表现差异的问题。
This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.