Score
Designs and writes precise, unambiguous documents that specify functional and non‑functional requirements, interfaces, component behavior, performance metrics, tolerances, test and acceptance criteria, and version/traceability information to guide development, manufacturing, or deployment. Analyzes and validates those specifications for completeness, consistency, feasibility, measurability, and standards compliance, and produces change‑controlled revisions and traceability artifacts to support verification and integration.
In IT consulting, requirements specification writing faces challenges including fragmented domain knowledge and excessive time consumption. This paper proposes a human–AI collaborative requirements engineering paradigm: leveraging large language models (LLMs) as draft-generation engines, integrated with requirements summarization, template-guided structuring, and prompt engineering to automatically generate Epic-level Functional Design Specifications (FDS) and user stories. Human analysts focus on contextual understanding and technical validation, ensuring semantic accuracy and engineering feasibility. Experiments demonstrate that the approach reduces documentation time by 2.3× on average and cuts human effort by ~40%. Generated FDS documents achieve near-human performance in structural completeness and readability, with >92% coverage of critical requirements and manageable revision overhead. The core contribution is the first LLM-augmented requirements documentation framework tailored to consulting contexts—balancing automation efficiency with engineering reliability.
In safety-critical aerospace software development, reconciling DO-178C compliance with agile iteration remains challenging due to inherent tensions between rigorous certification requirements and iterative, incremental practices. Method: This paper proposes a model-driven engineering (MDE) and metamodeling-based approach for continuous documentation generation. It establishes an extensible, automated documentation pipeline ensuring end-to-end traceability across requirements, design, and test evidence—enabling automatic generation of Requirements Traceability Matrices (RTMs), version-aware merge capabilities, and ternary audit tracing by role, time, and artifact. Contribution/Results: The method shifts certification from a monolithic, waterfall-style final activity to a verifiable, agile iterative process. Evaluated on a real avionics project, it achieves >90% automation rate for certification artifacts, 100% traceability completeness, and a 40% reduction in iteration cycle time—establishing the first industrial-grade paradigm for agile adoption in high-assurance domains.
Non-functional requirements (NFRs) are frequently missing or difficult to identify early in software engineering, particularly from functional requirements (FRs). Method: This paper proposes the first quality-attribute-driven, large language model (LLM)-assisted NFR generation framework. It integrates customized prompt engineering with a Deno pipeline and strictly aligns with the ISO/IEC 25010:2023 standard. The framework supports collaborative NFR generation across eight state-of-the-art LLMs (e.g., Gemini-1.5-Pro, Llama-3.3-70B). Contribution/Results: We conduct the first multi-LLM comparative evaluation, generating 1,593 NFRs from 34 FRs. Expert assessment yields average scores of 4.63/5.0 for NFR validity and 4.59/5.0 for attribute appropriateness, with 80.4% accuracy in quality-attribute classification—demonstrating the feasibility and practicality of automating high-quality NFR derivation in requirements engineering.
This work addresses the fragility of traditional requirements traceability in safety-critical software development, where reliance on external documentation often leads to silent breakdowns as code, requirements, and tests evolve independently. To overcome this, the paper proposes internalizing traceability as an intrinsic property of code structure by introducing language-native “Traceable” elements that enable compile-time verification of bidirectional links among requirements, implementations, and tests. The approach integrates code generation, metadata embedding, and build-time validation into the development workflow, providing proactive traceability assurance. When requirement changes disrupt traceability chains, the system automatically triggers warnings or build failures, thereby preventing traceability decay and significantly enhancing maintainability and reliability throughout software evolution.
Prior research lacks empirical insights into the actual adoption, quality effectiveness, and influencing factors of use case (UC) descriptions in industrial practice. Method: Drawing on large-scale industrial data from multinational enterprises (2020–2024), this study employs a mixed-methods approach—combining large-scale textual statistics, expert-based quality assessment, and regression/correlation analyses—to systematically examine UC practices. Contribution/Results: It provides the first large-scale empirical evidence revealing systematic deviations between real-world UC usage and textbook norms. Findings indicate that only a few features—such as solution orientation—significantly enhance development efficiency; UC template adoption exhibits a nonlinear relationship with quality; and UC quality exerts limited, highly context-dependent influence on downstream development—challenging conventional assumptions about UC quality efficacy. These results offer novel empirical evidence and theoretical directions for requirements engineering research and practice.
This study addresses the inefficiencies and impeded knowledge transfer arising from fragmented verification and validation (V&V) practices at the Jet Propulsion Laboratory (JPL). To overcome these challenges, this work proposes a unified V&V architecture grounded in human-centered design. By decoupling methodologies while maintaining a common attribute set, the architecture achieves bidirectional traceability through relational design and platform-independent SysML modeling. Furthermore, it establishes a comprehensive toolchain by integrating the Jama platform, modular templates, and digital thread technologies. This research effectively balances engineering rigor with agility, facilitating process automation, pattern reuse, and efficient cross-project collaboration. Ultimately, it provides a scalable and unified paradigm for the V&V of complex systems.
This study addresses the limitation of existing templates in specification-driven development, which fail to evaluate specification clarity and completeness. To overcome this, we propose EPIC, a framework grounded in the ISO/IEC/IEEE 29148 standard that conducts quantitative assessments of open-source repositories. By distilling an optimal specification taxonomy encompassing ten quality dimensions and forty practices, EPIC guides developers in clarifying expectations and bridging specification gaps. Empirical evaluations demonstrate that high-quality specifications reduce the proportion of bug-fixing commits to 11.8% and yield a fourfold increase in the median number of contributors. These findings indicate that adopting rigorous specification practices significantly enhances both collaborative efficiency and software quality in open-source projects.
研究通过采访技术作家探讨了软件文档审查过程及其挑战,识别了五个审查阶段,并揭示了组织和技术上的难题。
This study addresses the lack of realistic benchmarks and quality validation for AI agent-generated user documentation by constructing the first benchmark tailored to real-world maintenance scenarios. Leveraging documentation update tasks triggered by code changes across 292 open-source projects, it introduces a fine-grained, multi-dimensional automated evaluation metric incorporating multi-agent trajectory analysis, an abstention mechanism, and maintainer verification to systematically assess agents’ capabilities in updating or appropriately abandoning documentation. Experiments reveal that the best-performing agent achieves only 47.3 points, exposing critical deficiencies including neglecting the reader’s perspective, lacking evidential support, and overlooking impact scope. This work fills a significant benchmark gap in the field and provides essential empirical foundations for improving the quality of AI-generated documentation.
研究通过对比分析三种SBOM生成工具在JavaScript和Rust项目中的表现,揭示了因SBOM规范模糊导致的系统性差异问题,并建议未来需明确标准化规则以提高互操作性和合规性。