Score
Designs and evaluates the procedures, criteria, artifacts, and evidence used to formally accept a work product, including checklists, acceptance tests, traceability matrices, and authorization records. Builds signoff methodologies and performs verification activities that confirm requirements, compliance, and readiness for release.
This work addresses the challenge of reconciling task-level verification and regulatory traceability within high-velocity AI-assisted engineering workflows. The authors propose an “infinite loop” framework that integrates agile iteration with V-model validation, embedding independent verification and compliance auditing into every development cycle through a multi-agent AI architecture. The system automatically generates audit-ready documentation and incorporates critical human-in-the-loop approval gates. By natively embedding compliance capabilities into the development process, the approach achieves 100% requirement-level verification and enables trustworthy delivery with minimal human intervention. In a hardware-in-the-loop case study, the system attained full requirement pass rates with an average of only six human prompts per cycle, demonstrating a projected cost reduction of 10–50× compared to conventional methods.
In safety-critical aerospace software development, reconciling DO-178C compliance with agile iteration remains challenging due to inherent tensions between rigorous certification requirements and iterative, incremental practices. Method: This paper proposes a model-driven engineering (MDE) and metamodeling-based approach for continuous documentation generation. It establishes an extensible, automated documentation pipeline ensuring end-to-end traceability across requirements, design, and test evidence—enabling automatic generation of Requirements Traceability Matrices (RTMs), version-aware merge capabilities, and ternary audit tracing by role, time, and artifact. Contribution/Results: The method shifts certification from a monolithic, waterfall-style final activity to a verifiable, agile iterative process. Evaluated on a real avionics project, it achieves >90% automation rate for certification artifacts, 100% traceability completeness, and a 40% reduction in iteration cycle time—establishing the first industrial-grade paradigm for agile adoption in high-assurance domains.
This study addresses widespread compliance issues in GitHub Actions workflows, such as excessive permissions and weak secret management. It proposes the first documentation-driven compliance checking framework, which derives a 30-item checklist from official documentation and implements a hybrid auditing pipeline combining large language models (LLMs) with expert oversight. The authors automatically evaluate 95 real-world Java workflows using four open-source LLMs, employ GPT-5 as a conflict arbitrator, and integrate manual review into a multi-tiered validation system. Experimental results reveal an overall compliance rate of only 28%, with permission control as low as 4%. The proposed approach reduces manual verification effort by 81% while achieving 87% agreement with expert judgments, significantly enhancing audit efficiency and reproducibility.
This work addresses the problem of global inconsistency in multi-component intelligent agent releases, where local validation passes but cross-component relational integrity fails due to the absence of holistic consistency guarantees. To tackle this, we propose the Schema-SIP Relational Consistency (SIP-RC) framework—the first systematic approach to formally define and mitigate relational inconsistency faults in multi-component deployments. SIP-RC models release packages as graph structures and integrates schema documentation with product contract principles to enable cross-component relational verification. Key mechanisms include declarative–evidential linkage, decision authority scoping, provenance tracking of derived components, and byte-level consistency checks. Preliminary experiments demonstrate the feasibility of the proposed framework, offering a practical and actionable paradigm for ensuring relational consistency in intelligent agent releases.
Agile development methodologies struggle to comply with stringent airworthiness standards—such as DO-178C—in safety-critical aerospace software, due to inherent tensions between iterative flexibility and rigorous certification requirements. Method: This paper introduces CertiA360, an automated tool for end-to-end requirement traceability and compliance verification across agile iterations. Built upon DO-178C/DO-330, it features a certifiable architecture supporting dynamic requirement maturity assessment, change-driven verification and validation (V&V) closure, and bidirectional traceability. Contribution/Results: CertiA360 uniquely integrates lightweight agile practices—including user story mapping and incremental delivery—with high-assurance certification constraints, effectively bridging the gap between agility and regulatory rigidity. Empirical evaluation demonstrates a ~60% reduction in manual traceability effort, a 40% decrease in change-response cycle time, and full compliance with EASA/FAA tool qualification and process trustworthiness requirements.
This study addresses the inefficiencies and impeded knowledge transfer arising from fragmented verification and validation (V&V) practices at the Jet Propulsion Laboratory (JPL). To overcome these challenges, this work proposes a unified V&V architecture grounded in human-centered design. By decoupling methodologies while maintaining a common attribute set, the architecture achieves bidirectional traceability through relational design and platform-independent SysML modeling. Furthermore, it establishes a comprehensive toolchain by integrating the Jama platform, modular templates, and digital thread technologies. This research effectively balances engineering rigor with agility, facilitating process automation, pattern reuse, and efficient cross-project collaboration. Ultimately, it provides a scalable and unified paradigm for the V&V of complex systems.
本文提出Agile-V Assurance Spine,通过权威源配置文件、工件绑定、风险适当独立性和时效性等方法解决工程生命周期中对代理输出的正当行动问题。
This study addresses the lack of realistic benchmarks and quality validation for AI agent-generated user documentation by constructing the first benchmark tailored to real-world maintenance scenarios. Leveraging documentation update tasks triggered by code changes across 292 open-source projects, it introduces a fine-grained, multi-dimensional automated evaluation metric incorporating multi-agent trajectory analysis, an abstention mechanism, and maintainer verification to systematically assess agents’ capabilities in updating or appropriately abandoning documentation. Experiments reveal that the best-performing agent achieves only 47.3 points, exposing critical deficiencies including neglecting the reader’s perspective, lacking evidential support, and overlooking impact scope. This work fills a significant benchmark gap in the field and provides essential empirical foundations for improving the quality of AI-generated documentation.
研究通过采访技术作家探讨了软件文档审查过程及其挑战,识别了五个审查阶段,并揭示了组织和技术上的难题。
This study addresses the absence of concrete mapping mechanisms for implementing the EU AI Act within agile teams. Employing a Design Science Research methodology, this work proposes a novel framework that translates abstract regulatory requirements into actionable agile compliance guidelines. Through a traffic light taxonomy and expert interviews, an action catalog comprising twelve practices covering roles and risk management was constructed. The results demonstrate that these guidelines are both comprehensible and relevant, establishing that compliance should be integrated into existing agile activities rather than treated as a parallel process. Ultimately, this research bridges the gap in regulatory operationalization, providing a reusable methodological foundation that enables agile teams to achieve compliance without compromising iterative efficiency.