Score
Designs mappings from model components and learned features to formal theoretical constructs and builds methods that translate internal representations into testable, theory-aligned statements; analyzes post-hoc explanation procedures and evaluation metrics to test hypotheses and quantify trade-offs between explainability and model adaptability.
This work addresses the prevailing lack of systematic understanding of foundational formal theories in current AI compiler design, which hinders rigorous evaluation of the completeness and desirability of intermediate representations and compilation abstractions. For the first time, it systematically establishes precise correspondences between core mechanisms of MLIR—such as term rewriting systems, refinement calculi, and abstract interpretation—and classical formal theories. By grounding compiler abstractions in formal semantics, the paper clarifies the theoretical underpinnings of these constructs, articulates a precise notion of “design completeness,” and provides assessable criteria and guiding principles to navigate trade-offs between engineering pragmatism and theoretical ideals.
The absence of a mathematically formalized representation of design spaces renders design decisions heavily experience-dependent, hindering the development of automated support. Method: This paper introduces an orthogonal discretization model for design spaces—establishing the first structured spatial representation—and integrates, for the first time, large language model (LLM)-driven constraint generation with Monte Carlo tree search (MCTS) to enable autonomous, efficient exploration. It further develops a domain-adaptive instantiation engine that maps abstract design decisions to concrete implementations. Contribution/Results: The framework exhibits cross-domain transferability. Empirical evaluation on data article generation and chart visualization tasks demonstrates significant performance gains over baselines. User studies and expert interviews confirm its effectiveness, usability, and measurable improvement in design quality.
Modeling the structural representation of scientific methodology units and their cross-problem recombination mechanisms remains challenging, particularly in identifying common patterns among historically disruptive method combinations and discovering high-potential knowledge recombination pathways for novel problems. Method: We propose a novel framework comprising (1) contrastive learning to automatically extract structured representations of disruptive method combinations from multi-domain scientific literature, and (2) a reasoning-guided Monte Carlo search algorithm that integrates large language model (LLM)-based chain-of-thought reasoning with empirically derived historical innovation patterns to enable interpretable, goal-directed knowledge recombination. Contribution/Results: Empirical evaluation across physics, biology, and artificial intelligence demonstrates that our framework accurately identifies method combinations with high disruptive potential and significantly advances the modeling and predictive capability of scientific innovation dynamics—achieving improved fidelity in capturing structural evolution and recombination efficacy in scientific discovery processes.
Large language models (LLMs) struggle to consistently preserve object-oriented design intent in software design synthesis, exhibiting output non-determinism and sensitivity to prompting. This work introduces the first benchmark for evaluating object-oriented design intent fidelity and systematically assesses the reliability of ChatGPT-4o-mini, Claude 3.5 Sonnet, and Gemini 2.5 Flash in generating UML class diagrams under standard prompting, rule injection, and a novel preference-aligned few-shot prompting strategy. Experimental results demonstrate that preference alignment substantially improves adherence to design intent but fails to eliminate non-determinism; moreover, inherent model behaviors significantly influence reliability. This study provides the first empirical analysis of LLM-based design synthesis stability through the lenses of non-determinism, prompt sensitivity, and methodological scaffolding, establishing a new paradigm for dependable AI-assisted software design.
Empirical evidence on the impact of Model-Driven Engineering (MDE) on software quality is fragmented and lacks systematic integration. Method: This paper conducts the first tertiary study dedicated to MDE quality research, systematically analyzing 22 published systematic literature reviews and mapping studies. It establishes a three-tier analytical framework to characterize research distribution, evidential strength, and methodological maturity in the MDE–quality domain. Results: Maintainability is the most studied quality attribute; however, among 83 identified research questions, 80 focus solely on conceptual or syntactic model-to-code mappings, with few conducting empirical comparisons. Crucially, MDE’s actual impact on quality in industrial development contexts remains markedly under-investigated. The study exposes a structural bias toward “re-modeling over validation” in current research and identifies critical gaps requiring urgent attention: rigorous experimental design, industry-based empirical validation, and multi-attribute quality assessment frameworks.
This work addresses the inefficiencies inherent in manual construction of reference and influence models within Design Research Methodology (DRM), including poor readability, difficulty in modification, and challenges in tracing evidential support. To overcome these limitations, the paper introduces DREAMS, a novel modeling environment that, for the first time, articulates DRM-specific modeling support requirements. DREAMS incorporates typed causal modeling and symbolic relationship representation, directly anchoring hypotheses, empirical inputs, and literature citations to causal links. It further integrates interactive layout optimization and efficient retrieval mechanisms. Preliminary user evaluations demonstrate that DREAMS significantly reduces model creation and revision time, minimizes node reordering and edge crossings, and enhances both evidential traceability and overall model maintainability.
This work addresses the lack of a unified theoretical foundation in interpretable machine learning, which has led to fragmented methodologies and inconsistent evaluation criteria. By introducing Lagrangian mechanics into this domain for the first time, the paper proposes a general theoretical framework grounded in user-oriented interpretability. Through systematic analysis of symmetries and constraints, the approach derives optimal interpretable models by minimizing a suitably defined Lagrangian. This deductive methodology not only unifies existing techniques under a coherent theoretical umbrella but also reveals novel research directions. It has successfully informed the design of core programming interfaces, mitigated limitations of current methods, and established a rigorous theoretical basis for interpretability education and interdisciplinary integration.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
This study addresses the limited semantic transparency and poor comprehensibility of existing conceptual models, which stem from their reliance on low-level syntactic constructs to represent domain abstractions, thereby hindering effective system design and stakeholder communication. To overcome this, the paper proposes a language-agnostic abstract symbol engineering approach that identifies, formalizes, visualizes, and validates recurring syntactic configuration patterns, replacing them with high-level, semantically transparent abstract symbols. The method is instantiated as the DeCleaR extension to Dynamic Condition Response (DCR) graphs. Empirical evaluation demonstrates that DeCleaR significantly enhances perceived model quality, pragmatic quality, and user preference compared to standard DCR graphs.
This study addresses the limited external validity of software engineering experiments, often stemming from unrepresentative samples. It pioneers the systematic application of causal inference–based transportability methods in this domain, integrating experimental and observational data to develop tailored implementation pathways and practical guidelines. The proposed approach is validated through simulation studies and offers actionable strategies for generalizing findings from common yet constrained settings—such as using students as proxies for professional developers—to broader target populations. By explicitly modeling the mechanisms underlying population differences, the method significantly enhances the practical applicability and reliability of experimental results across diverse real-world contexts.