Score
Designs and runs structured architecture and design review processes for software and system technical designs, analyzing design rationale, interfaces, constraints, tradeoffs, quality attributes, and compliance with requirements to identify risks and recommended mitigations. Produces and maintains review artifacts such as checklists, reports, action items, and leads review meetings to track issue resolution and enforce design standards.
This study addresses the low efficiency and high cost of manual review for software design documentation by proposing an LLM-based automation framework for multi-perspective review. Methodologically, it systematically categorizes eleven review dimensions and identifies those amenable to automation using general-purpose large language models (e.g., GPT); for hybrid semi-structured documents containing tables, it integrates prompt engineering with context-aware modeling to enhance the model’s comprehension of design logic and cross-document consistency. The key contribution is overcoming current LLM limitations in interpreting structured information, enabling— for the first time—the automated detection of cross-document design inconsistencies across multiple perspectives. Experimental evaluation on real-world industrial documentation demonstrates that the approach accurately identifies inconsistency issues, achieving an accuracy level meeting industrial deployment requirements.
This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.
In software design, paradigm-implied semantic expectations—such as data abstraction consistency and feedback-control closed-loop behavior—are often left implicit, leading to design deviations and verification challenges. To address this, we introduce the concept of *design obligations*: explicit, logically formalizable, and verifiable specifications that codify such implicit constraints inherent to design paradigms. Leveraging formal modeling and paradigm semantics analysis, we establish two obligation frameworks—one for data-abstraction-based systems and another for feedback-driven adaptive systems—precisely capturing their core semantic requirements. We demonstrate that common design flaws stem from obligation violations and show how these obligations enable rigorous compliance verification and pedagogical application. This work bridges the semantic gap between design intent and implementation, providing both theoretical foundations and a methodological framework for paradigm-driven design assurance.
Prior work lacks empirical characterization of problem-solving processes in software development. Method: Integrating grounded theory coding, sequential pattern mining, and multidimensional statistical analysis on 356 Mozilla Firefox issue reports, this study extracts fine-grained, reusable problem-solving process patterns from collaborative textual artifacts. Contribution/Results: We identify 47 empirically grounded process patterns—challenging the traditional linear assumption by revealing pervasive nonlinearity: 73% of fixes involve iterative backtracking or parallel activities. The resulting process landscape and pattern catalog systematically characterize distributional regularities across issue types, defect categories, and repair durations. This advances understanding of real-world engineering complexity and provides an evidence-based foundation for process optimization, collaborative tool design, and developer support.
In organizations lacking a formal architect role, software architecture decisions are frequently made by practitioners without such titles. This study employs a mixed-methods approach, combining a survey of 54 practitioners with in-depth interviews of 7 participants, to investigate who actually makes architectural decisions and under what contextual conditions. The findings reveal that informal architects are extensively involved in critical architectural choices, while formally designated architect roles are primarily deemed necessary in large enterprises or complex teams. These results challenge the prevailing research paradigm that centers on formal architects, instead illuminating how architectural responsibilities are distributed across real-world development environments and highlighting their strong dependence on organizational context.
This work addresses a critical limitation in current AI-based code review systems: in the absence of executable specifications, they often fall into structural loops and exhibit correlated errors due to shared training distributions between generation and review models, making it difficult to verify whether code aligns with true intent. The paper proposes a three-tiered architecture—“specification-first, deterministic verification, AI review of residual issues”—and provides the first systematic demonstration that executable specifications can transform code review from a complex domain into a complex yet solvable one. It further clarifies that AI should focus specifically on structural and architectural flaws beyond specification coverage. Through deterministic verification, cross-model review, and targeted defect injection experiments, the study confirms that both intra- and inter-family large models exhibit error correlation without specification guidance, whereas a specification-driven architecture effectively isolates the AI review boundary and significantly enhances reliability.
This work proposes a systematic approach to derive task effectiveness requirements in the absence of explicit user needs. The method deconstructs task intent into context, functionality, constraints, critical dimensions, performance attributes, and architectural solutions, and introduces a task complexity factor to quantify the impact of external challenges and technology maturity. By integrating Best-Worst Scaling, it prioritizes critical dimensions based on stakeholder judgments. Through task decomposition modeling and quantitative complexity analysis, the framework supports integration with UAF/SysML artifacts and establishes a traceable mechanism for generating Tier 1 and Tier 2 requirements. The approach is validated using a close air support mission case study, effectively addressing a critical gap in requirements engineering when clear initial inputs are unavailable.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).