Score
Designs, authors, and maintains specification artifacts that define system- or product-level requirements, functional behavior, interfaces, component responsibilities, architectural constraints, and acceptance criteria. Develops formal or model-based specifications and performs specification analysis (completeness, consistency, testability, and traceability) to guide implementation, verification, and integration.
In IT consulting, requirements specification writing faces challenges including fragmented domain knowledge and excessive time consumption. This paper proposes a human–AI collaborative requirements engineering paradigm: leveraging large language models (LLMs) as draft-generation engines, integrated with requirements summarization, template-guided structuring, and prompt engineering to automatically generate Epic-level Functional Design Specifications (FDS) and user stories. Human analysts focus on contextual understanding and technical validation, ensuring semantic accuracy and engineering feasibility. Experiments demonstrate that the approach reduces documentation time by 2.3× on average and cuts human effort by ~40%. Generated FDS documents achieve near-human performance in structural completeness and readability, with >92% coverage of critical requirements and manageable revision overhead. The core contribution is the first LLM-augmented requirements documentation framework tailored to consulting contexts—balancing automation efficiency with engineering reliability.
This work addresses a critical gap in the evaluation of software engineering agents, which has predominantly focused on code implementation while neglecting their ability to detect and correct defects in requirements specifications—such as omissions, ambiguities, and inconsistencies. We propose the first evaluation framework centered on specification-level reasoning, constructing a benchmark based on the RFC (Request for Comments) processes of open-source projects. The framework requires agents to systematically identify design flaws by synthesizing initial proposals, code repositories, and historical discussions. Evaluations across five repositories, including Kubernetes and React, reveal that even the best-performing model (GPT-5.4) achieves only 44.4% accuracy, highlighting a significant limitation in current agents’ capacity for requirement analysis and design review without execution feedback. This study thus fills a crucial void in assessing agent capabilities at the specification stage.
This paper addresses two key challenges in formal modeling of space-system requirements: high ambiguity in natural-language specifications and weak cross-paradigm traceability. Building upon NASA’s FRET toolchain, it presents the first systematic approach for automated, bidirectional translation from natural-language requirements to multi-paradigm formal specifications—namely Linear Temporal Logic (LTL), Architecture Analysis & Design Language (AADL), and Systems Modeling Language (SysML)—along with rigorous bidirectional traceability verification. Methodologically, it introduces a unified framework supporting requirement–specification bidirectional mapping, integrating temporal-logic verification with semantic alignment across architectural and modeling languages. Contributions include: (1) 100% requirement coverage and fully structured traceability chains across four real-world space-system case studies; (2) empirical characterization of expressive boundaries and interoperability pathways among formal paradigms; and (3) significant improvements in ambiguity detection and specification consistency verification efficiency, establishing a reusable methodology for high-assurance space-system requirements engineering.
This work addresses the challenge of reliably conveying intent, requirements, and constraints in human–AI–tool collaborative software development by proposing a specification-centric Bosque API (BAPI) ecosystem. The system introduces a highly expressive specification language that, for the first time, enables cross-language interoperability, automated test generation, formal verification, and execution sandboxing across the entire API lifecycle—from requirement definition and implementation to invocation and validation. By providing end-to-end specification guarantees, BAPI significantly enhances system correctness, security, and the efficiency of human–AI collaboration, offering a novel infrastructure for software development in the era of AI agents.
This study addresses the disconnect between comprehension and execution in large language model (LLM) agents, which frequently leads to falsely reported task completion. To mitigate this issue, we propose SpecHarness, a framework that compiles visible specifications into source-linked obligations, decoupling agent proposals from authoritative state. Through runtime-verifiable mediation mechanisms and versioned state management, SpecHarness enables agent-independent compliance verification. Experimental results demonstrate that the proposed approach effectively bridges cognitive gaps and significantly reduces false completion rates, ensuring that tasks strictly adhere to external specifications during execution. Ultimately, this work provides a reliable, architecture-level solution for governing LLM agent behavior.
This study addresses the challenge of transforming stakeholder requirements into product requirements in software-driven automotive systems. Leveraging a dataset of 8,082 stakeholder requirements and 5,870 product requirements provided by Infineon, the research employs a hybrid methodology integrating structural statistics, decision modeling, traceability mining, textual analysis, and hardware-software linkage to systematically analyze the requirement refinement process. It reveals, for the first time, that requirement complexity primarily stems from ambiguous architectural scope and missing contextual information rather than linguistic redundancy. The work establishes a classification framework for mapping stakeholder to product requirements, identifies systematic differences across abstraction levels, and proposes key improvements in requirement validation, deviation management, and contextual tooling to support efficient and reusable automotive development.
本文提出了一种形式化的语义块模型和执行评判基准来独立评估规范质量,通过结构化表示和机器可验证条件解决规范确定性问题。
This study addresses the limitation of existing templates in specification-driven development, which fail to evaluate specification clarity and completeness. To overcome this, we propose EPIC, a framework grounded in the ISO/IEC/IEEE 29148 standard that conducts quantitative assessments of open-source repositories. By distilling an optimal specification taxonomy encompassing ten quality dimensions and forty practices, EPIC guides developers in clarifying expectations and bridging specification gaps. Empirical evaluations demonstrate that high-quality specifications reduce the proportion of bug-fixing commits to 11.8% and yield a fourfold increase in the median number of contributors. These findings indicate that adopting rigorous specification practices significantly enhances both collaborative efficiency and software quality in open-source projects.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
This work addresses the high cost of manually writing formal specifications and the limitations of existing large language model (LLM)-based approaches that require white-box access to source code, thereby posing intellectual property and deployment constraints. The authors propose a black-box-driven method that leverages only test code and dynamic execution traces to generate candidate Java Modeling Language (JML) specifications via an LLM. These candidates are locally validated using bounded model checking, and an iterative feedback loop refines them based on verification outcomes. This approach is the first to enable fully automated formal specification generation without any access to the program’s internal structure. Evaluated on the SpecGenBench benchmark, it demonstrates that test-derived information effectively guides specification synthesis, while also highlighting critical challenges in checker compatibility and diagnostic feedback, substantially enhancing industrial applicability.