Score
Designs, builds, or analyses systems and workflows that disambiguate written requirements by retrieving relevant evidence and contextual documents, generating candidate clarified requirement formulations, and presenting ranked alternatives. The work includes constructing retrieval pipelines for clause-level evidence, modules to produce and score candidate disambiguations by relevance and clarity, and mechanisms for analyst-in-the-loop validation and selection.
This study addresses pragmatic ambiguity in natural language requirements, which often arises from discrepancies in stakeholders’ domain knowledge and contextual understanding. The authors propose a novel approach that integrates a multi-level domain knowledge base with Retrieval-Augmented Generation (RAG) to simulate stakeholders at beginner, intermediate, and expert proficiency levels. This framework systematically identifies divergent interpretations of requirements and generates candidate disambiguated formulations grounded in expert knowledge, which are subsequently validated by requirements analysts. Evaluated on the PURE dataset, the method effectively detects pragmatic ambiguities, with GPT-4o-mini achieving a recall and F2 score of 0.75. Human evaluation further demonstrates that the generated disambiguated requirements are highly effective: GPT-4o-mini excels in relevance, while Mistral-7B leads in clarity and consistency.
Ambiguity in natural-language requirements frequently leads large language models (LLMs) to generate incorrect code. To address this, we propose SpecFix—a fully automated ambiguity resolution method that requires no LLM metacognitive capabilities. SpecFix first models the distribution of program interpretations induced by an LLM’s responses to the original requirement, using program testing and automated repair. It then infers a revised, minimal, and verifiable specification by reversely analyzing distributional shifts and iteratively contracting the requirement space under logical constraints. Crucially, SpecFix decouples ambiguity resolution into two novel phases: (1) modeling the program-interpretation distribution and (2) constraint-driven, contraction-based specification inference—thereby eliminating reliance on self-reflection. Evaluated on HumanEval+ and MBPP+, SpecFix improves Pass@1 by 4.3% on average across GPT-4o, DeepSeek-V3, and Qwen2.5-Coder, and boosts majority-vote solution rates by 3.4%.
This work addresses the error-prone and labor-intensive process of manually translating regulatory texts such as the GDPR and the EU AI Act into actionable software requirements. The authors propose Reg2Req, the first end-to-end automated pipeline that leverages natural language processing to identify regulatory provisions, generate system-agnostic software requirements accompanied by plain-language explanations, and establish traceability links. The approach supports requirement classification, use case seed generation, and cross-reference analysis, achieving macro-averaged F1 scores of 0.82 on the GDPR and 0.78 on the EU AI Act. A user study demonstrates that the generated plain-language explanations significantly enhance users’ comprehension and confidence in taking compliance actions (p < 0.001), with all participants expressing willingness to adopt the output as a starting point for compliance efforts.
In enterprise RAG scenarios, large language models (LLMs) struggle with domain-ambiguous queries and often introduce lexical or semantic noise. To address this, we propose VERDICT—a novel framework featuring a Verified-Diversification mechanism that enables collaborative, feedback-driven interaction between retriever and generator. Unlike conventional cascaded “generate–retrieve–filter” pipelines, VERDICT integrates differentiable early-stage verification with diverse explanation generation in a closed loop, mitigating error propagation. Our method unifies dynamic explanation generation, retrieval feedback modeling, and consistency consolidation. Evaluated on the ASQA benchmark, VERDICT achieves an average 23% improvement in grounded F1 over strong baselines. Moreover, it demonstrates robust generalization across multiple mainstream LLM backbones. VERDICT establishes a new paradigm for ambiguity resolution in RAG—robust, efficient, and scalable—while preserving interpretability and grounding fidelity.
To address challenges in requirements engineering—including difficulty in identifying relationships among natural language requirements, high manual annotation costs, and poor domain adaptability—this paper proposes an NLP-driven, systematic relation extraction framework. It is the first to integrate a requirements relationship ontology with multi-paradigm NLP techniques: dependency parsing, semantic role labeling, named entity recognition, BERT-based supervised fine-tuning, and retrieval-augmented methods. A unified classification-based evaluation framework is established to clarify core challenges and evolutionary pathways. The framework supports major requirement relations (e.g., *refines*, *conflicts*) and enables reusable, extensible relation modeling. Experimental results demonstrate significant improvements in automation capability and accuracy for large-scale adaptive requirements management systems, thereby strengthening requirements evolution analysis and consistency verification.
This study addresses the prevalent ambiguity, inconsistency, and incompleteness in articulating explainability requirements for AI systems due to a lack of standardized specifications. Through a structured literature review and interviews with developers, the authors identify a set of explainability quality attributes, which are then refined via a large-scale survey of practitioners into ten core attributes. For the first time, these attributes are translated into a prioritized, actionable guideline for writing explainability requirements. Building on this foundation, the authors design a lightweight, iterative requirements engineering workflow augmented by a large language model to assist in requirement generation. An accompanying web-based tool reduces average requirement drafting time by 23.5%, and user evaluations indicate that the generated requirements match or slightly exceed manually written ones in terms of implementability and textual quality.
Software requirements are often implicit in stakeholder interviews, making them difficult to capture explicitly yet critically important for system design. This work proposes LENS, a novel approach that leverages context-aware large language models (LLMs) to jointly extract explicit requirements and infer implicit ones from interview transcripts, while incorporating organizational context to generate traceable user stories. LENS enables unified modeling and traceability of both explicit and implicit requirements. Evaluated on 12 interview transcripts from the cybersecurity domain, the method achieves an F1 score of 84.4% in explicit requirement extraction, and 75% of the inferred implicit requirements were rated by domain experts as practically valuable, demonstrating its potential to support automation and reduce manual analysis effort.
This work addresses the limited ability of large language models (LLMs) to handle ambiguous requirements in code generation, a challenge exacerbated by the lack of systematic evaluation of their clarification capabilities. We introduce ClarifyCodeBench, the first interactive benchmark grounded in real-world programming tasks, which features human-annotated ambiguity types, clarification questions, and reference answers to enable structured assessment. To quantify clarification effectiveness, we propose two novel metrics: Turn-discounted Key Question Rate and Optimal Round Adherence. Our comprehensive evaluation reveals that strong code generation performance does not necessarily entail effective clarification; reasoning augmentation offers minimal gains in ambiguity detection; and model performance degrades substantially in multi-ambiguity scenarios—evidence of a decoupling between code generation and clarification abilities.
This work addresses the challenge engineering teams face in explicitly identifying security requirements from unstructured to-do items in regulated domains. The authors propose a natural language processing–based to-do enhancement system that integrates a high-recall security relevance classifier with a four-stage retrieval-augmented generation (RAG) pipeline to automatically detect security-related entries and link them to compliance requirements. The study introduces the first publicly released dataset of security to-do items annotated by domain experts. The classifier achieves an in-distribution F2 score of 0.774 and demonstrates robust generalization, attaining an average zero-shot G-measure of approximately 0.65 across five benchmarks. In expert evaluation, 12 out of 24 regulatory clauses retrieved by the RAG pipeline received scores of at least 4 out of 5, confirming the method’s practical feasibility in industrial settings.