systematic model analysis

Designs and implements systematic processes, tools, and artifacts (e.g., evaluation protocols, benchmarks, catalogs, and reproducible workflows) to catalog, evaluate, compare, and reproduce models, datasets, and findings. Analyzes and synthesizes model behaviors and results across methods and levels of analysis to produce cumulative, reproducible knowledge and discipline-level standards for model assessment.

systematicmodelanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.

behavioral propertiesinterface propertiesmodel verification

The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources

Jun 24, 2024
SL
Shayne Longpre
🏛️ MIT | EleutherAI | UCSB | UW | Princeton University | Stanford University | Harvard University | Allen Institute for AI | Google DeepMind | ML Commons | HuggingFace | University College London

This paper addresses the systemic absence of responsible practices in foundational model development by introducing the first comprehensive, multimodal resource guide covering text, vision, and speech modalities. Through systematic literature review, cross-modal taxonomy construction, and tool-to-capability mapping, it identifies four critical structural gaps: (1) scarcity of multimodal and multilingual tooling; (2) weak capabilities in data curation and safety evaluation; (3) insufficient system-level monitoring and reproducibility infrastructure; and (4) lack of environmental impact assessment and release governance frameworks. The project delivers a curated practice inventory comprising 250+ open-source tools and resources spanning data governance, training optimization, safety auditing, carbon footprint analysis, and responsible deployment. Empirically grounded, the findings inform policy formulation, tool development, and standardization efforts—advancing AI development from heuristic practice toward a verifiable, auditable, and sustainable engineering paradigm.

Multimodal and multilingual analysis gapsResponsible foundation model developmentTools for ethical AI practices

Quality in model-driven engineering: a tertiary study

Jun 23, 2016
MG
M. Goulão
🏛️ Universidade Nova de Lisboa | University of Maribor

Empirical evidence on the impact of Model-Driven Engineering (MDE) on software quality is fragmented and lacks systematic integration. Method: This paper conducts the first tertiary study dedicated to MDE quality research, systematically analyzing 22 published systematic literature reviews and mapping studies. It establishes a three-tier analytical framework to characterize research distribution, evidential strength, and methodological maturity in the MDE–quality domain. Results: Maintainability is the most studied quality attribute; however, among 83 identified research questions, 80 focus solely on conceptual or syntactic model-to-code mappings, with few conducting empirical comparisons. Crucially, MDE’s actual impact on quality in industrial development contexts remains markedly under-investigated. The study exposes a structural bias toward “re-modeling over validation” in current research and identifies critical gaps requiring urgent attention: rigorous experimental design, industry-based empirical validation, and multi-attribute quality assessment frameworks.

Aggregating consolidated findings on quality impact in model-driven engineeringAnalyzing software quality attributes most affected by MDE approachesIdentifying under-explored research areas needing further empirical validation

Latest Papers

What's happening recently
View more

This study addresses the challenge of quantifying and comparing large language model (LLM) behaviors across vendors in a standardized, cost-effective manner. We propose a simple, inexpensive, and reproducible framework for investigating model behavior by applying a fixed set of public stimuli across a cross-vendor panel of models. The framework innovatively integrates three complementary evaluation methods—exact matching, LLM-judge codebooks, and instrumented environments—to enable scalable behavioral tracking at minimal cost. Experiments reveal lexical convergence among models, evolving robustness to suffix-based prompts, divergences in stance adherence, and patterns of documentation non-compliance exhibited by coding agents. Collectively, this work establishes a systematic evaluation paradigm for tracking the behavioral evolution of large language models.

behavioral measurementcross-vendor comparisonevaluation

This study addresses the absence of benchmark datasets for Modelica, which has hindered empirical research on model evolution. We propose ModBench, an automated pipeline that establishes a novel paradigm for generating model snapshot benchmarks directly from code repositories by mining Git history, filtering commits, extracting simulatable classes, and normalizing representations. Applying this approach to the Modelica Standard Library, we constructed a comprehensive dataset comprising 85,562 class snapshots spanning all versions since v3, complete with API access and traceability links. This work fills a critical data gap in the domain, providing essential infrastructure to support research in model evolution analysis, compiler testing, and automated program repair.

benchmark datasetscyber-physical systemsmodel evolution

This study addresses the lack of empirical evidence in data quality management for AI systems, where traditional perspectives struggle with model attribution and compliance challenges. Employing reflexive thematic analysis through in-depth interviews with 16 practitioners, this work examines the engineering and organizational dimensions of data quality in AI-driven systems, revealing emergent characteristics including traceability, circularity, and legitimacy. It introduces a novel conceptual framework termed “lifecycle assurance” that integrates fragmented machine learning research agendas and establishes evidence-generation mechanisms supporting specific AI claims. Furthermore, the study identifies six overarching themes and five trust-influencing conditions, offering practice-based, engineering-oriented guidance for managing data quality in AI systems.

AI-driven systemsdata qualityfoundation models

Hot Scholars

JL

Jiayi Liu

Meta Platforms
Data ScienceMachine LearningPhysicsCosmology
SZ

Siyuan Zhang

Tsinghua University
large language modelAI safetyreinforcement learning
KC

Kecheng Chen

PhD student at EE, City University of Hong Kong
Transfer LearningAI for HealthcareSignal Processing
TP

Tianyu Pang

Senior Research Scientist, Sea AI Lab
Machine LearningGenerative ModelsTrustworthy AI