Score
Designs and implements systematic processes, tools, and artifacts (e.g., evaluation protocols, benchmarks, catalogs, and reproducible workflows) to catalog, evaluate, compare, and reproduce models, datasets, and findings. Analyzes and synthesizes model behaviors and results across methods and levels of analysis to produce cumulative, reproducible knowledge and discipline-level standards for model assessment.
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
This paper addresses the systemic absence of responsible practices in foundational model development by introducing the first comprehensive, multimodal resource guide covering text, vision, and speech modalities. Through systematic literature review, cross-modal taxonomy construction, and tool-to-capability mapping, it identifies four critical structural gaps: (1) scarcity of multimodal and multilingual tooling; (2) weak capabilities in data curation and safety evaluation; (3) insufficient system-level monitoring and reproducibility infrastructure; and (4) lack of environmental impact assessment and release governance frameworks. The project delivers a curated practice inventory comprising 250+ open-source tools and resources spanning data governance, training optimization, safety auditing, carbon footprint analysis, and responsible deployment. Empirically grounded, the findings inform policy formulation, tool development, and standardization efforts—advancing AI development from heuristic practice toward a verifiable, auditable, and sustainable engineering paradigm.
Empirical evidence on the impact of Model-Driven Engineering (MDE) on software quality is fragmented and lacks systematic integration. Method: This paper conducts the first tertiary study dedicated to MDE quality research, systematically analyzing 22 published systematic literature reviews and mapping studies. It establishes a three-tier analytical framework to characterize research distribution, evidential strength, and methodological maturity in the MDE–quality domain. Results: Maintainability is the most studied quality attribute; however, among 83 identified research questions, 80 focus solely on conceptual or syntactic model-to-code mappings, with few conducting empirical comparisons. Crucially, MDE’s actual impact on quality in industrial development contexts remains markedly under-investigated. The study exposes a structural bias toward “re-modeling over validation” in current research and identifies critical gaps requiring urgent attention: rigorous experimental design, industry-based empirical validation, and multi-attribute quality assessment frameworks.
This study addresses the challenge of quantifying and comparing large language model (LLM) behaviors across vendors in a standardized, cost-effective manner. We propose a simple, inexpensive, and reproducible framework for investigating model behavior by applying a fixed set of public stimuli across a cross-vendor panel of models. The framework innovatively integrates three complementary evaluation methods—exact matching, LLM-judge codebooks, and instrumented environments—to enable scalable behavioral tracking at minimal cost. Experiments reveal lexical convergence among models, evolving robustness to suffix-based prompts, divergences in stance adherence, and patterns of documentation non-compliance exhibited by coding agents. Collectively, this work establishes a systematic evaluation paradigm for tracking the behavioral evolution of large language models.
研究解决数据仓库中缺乏变量级元数据的问题,通过提出一种与DDI-CDI模型兼容的元数据应用配置文件方法来丰富元数据并进行模型-数据一致性检查。
This study addresses the absence of benchmark datasets for Modelica, which has hindered empirical research on model evolution. We propose ModBench, an automated pipeline that establishes a novel paradigm for generating model snapshot benchmarks directly from code repositories by mining Git history, filtering commits, extracting simulatable classes, and normalizing representations. Applying this approach to the Modelica Standard Library, we constructed a comprehensive dataset comprising 85,562 class snapshots spanning all versions since v3, complete with API access and traceability links. This work fills a critical data gap in the domain, providing essential infrastructure to support research in model evolution analysis, compiler testing, and automated program repair.
This study addresses the lack of empirical evidence in data quality management for AI systems, where traditional perspectives struggle with model attribution and compliance challenges. Employing reflexive thematic analysis through in-depth interviews with 16 practitioners, this work examines the engineering and organizational dimensions of data quality in AI-driven systems, revealing emergent characteristics including traceability, circularity, and legitimacy. It introduces a novel conceptual framework termed “lifecycle assurance” that integrates fragmented machine learning research agendas and establishes evidence-generation mechanisms supporting specific AI claims. Furthermore, the study identifies six overarching themes and five trust-influencing conditions, offering practice-based, engineering-oriented guidance for managing data quality in AI systems.
为解决AI模型研究发现散乱问题,提出Modelpedia框架,利用LLM自动提取并整理论文中的模型发现,形成可搜索的公共目录。