Score
Designs and manages iterative development and release cycles for a product, including hypotheses, prototypes, experiments, feature releases, and success metrics to validate changes. Builds and analyzes feedback loops and telemetry, prioritizes backlog and implementation based on measured outcomes, and refines product scope, user experience, and performance across successive versions.
本文提出一种生命周期感知的框架,结合软件质量评估与大语言模型代码优化,以解决科研软件因开发者缺乏软件工程经验导致的质量问题。
This work addresses the challenge of balancing software quality, testability, and maintainability under rapid iteration and frequent requirement changes. It proposes Algorithm-Driven Development (ADD), a novel approach that unifies requirements specification and technical design by using algorithm flowcharts as a single, coherent artifact. This integration enables end-to-end modeling of requirements, architecture, and testing. Leveraging this model, the system automatically generates high-coverage acceptance tests and incorporates continuous integration with code coverage feedback. Industrial adoption at Dassault Systèmes demonstrates that ADD achieves over 95% code coverage, substantially reduces defect density, and ensures a stable delivery cadence, outperforming conventional test-driven development and test-after approaches.
This study investigates the mechanisms through which iterative and sequential workflows influence innovative behavior and performance. Through three controlled laboratory experiments integrating behavioral observation and performance analysis, the research systematically compares the effectiveness of these workflow types across diverse innovation tasks. It provides the first empirical evidence that iterative processes enhance performance not only in idea generation but also in non-creative tasks, primarily by increasing task-switching frequency and thereby broadening the search space for solutions. The study identifies three critical boundary conditions—tasks requiring extensive exploration, high inter-component coordination, and moderate time pressure—under which iterative workflows significantly boost innovation performance. However, this advantage diminishes as task characteristics shift or over time, offering nuanced insights into the dynamics of innovation search processes.
Existing research lacks systematic methods to assess how requirements engineering (RE) impacts downstream development activities, hindering RE process optimization. Method: This paper proposes the first fitness-for-purpose RE impact assessment model, integrating a systematic literature review with multi-source empirical data to identify and structure 24 downstream development activities affected by requirements and 16 quantifiable attributes. Contribution/Results: The model bridges two critical gaps in requirements quality assessment—namely, the “activity dimension” and “measurability of impact”—by enabling empirical analysis of how specific requirements artifacts and processes concretely influence development practices. It provides a theoretically grounded framework and evidence-based decision support for precise, targeted optimization of the RE phase.
Inconsistent definitions of “feature” across software engineering domains—particularly requirements engineering (RE) and software product lines (SPL)—impede communication, trigger rework, and reduce cross-team collaboration efficiency. Method: We conducted an empirical study across 27 mainstream open-source projects, integrating repository mining, branch behavior analysis, qualitative coding, and pattern induction to derive a data-driven, cross-disciplinary definition of feature. Contribution/Results: This work introduces the first empirically grounded, unified feature definition framework bridging RE and SPL. It identifies recurring collaboration patterns and critical bottlenecks in feature description, implementation, and management, and proposes a roadmap linking academic theory with industrial practice. The findings yield actionable guidelines for project planning, resource allocation, and inter-team coordination, advancing feature conceptual standardization and engineering practice optimization.
This study addresses the lack of systematic understanding regarding the impact of repair loop iteration counts in large language model (LLM)-based software engineering tasks, where prior work often relies on arbitrarily defined repair budgets. Through a cross-task (code generation, test generation, code translation) and cross-model empirical analysis, this work reveals—for the first time—a pronounced diminishing marginal returns phenomenon in iterative repair: performance gains are concentrated within the first 3–4 iterations, with negligible improvements thereafter. The findings underscore that the design of the repair workflow and feedback mechanisms exerts a far greater influence on repair efficacy than the choice of LLM itself. The authors advocate for treating repair budget as a critical experimental variable to ensure reliable, computationally efficient, and reproducible evaluation outcomes.
This study addresses the challenges of requirement drift, perceptual deficits, and accountability ambiguity in coding agent iterations by proposing a Human-Agent-Virtual User engineering closed-loop framework. Leveraging multi-agent collaboration, version binding, and virtual user simulation, this approach enables end-to-end traceability and intent verification spanning from requirement confirmation to automated testing. The method establishes an auditable development lifecycle that effectively mitigates requirement drift while ensuring human oversight of final releases. Consequently, it achieves accountable agent-based application delivery and provides a reliable human-AI collaboration paradigm for complex software development.
Existing approaches to requirement prioritization often overlook the semantic interdependencies among requirements, thereby compromising prioritization effectiveness. This work addresses this limitation by introducing requirement interconnectedness into user feedback–driven prioritization for the first time, proposing a dependency-aware search-based optimization framework. The method first applies natural language processing to cluster app store feedback into semantically coherent requirement groups and then automatically infers “requires”-type dependencies among these groups. These dependencies are explicitly integrated into a search algorithm to guide the optimization of requirement priorities. Evaluated on 94 real-world instances across four software systems, the proposed approach significantly outperforms ReFeed, demonstrating that explicitly modeling requirement interconnections effectively enhances both prioritization accuracy and release planning quality.
This work addresses the fragmentation among requirements, testing, and production phases in software engineering, which hinders continuous quality improvement due to the absence of a production-feedback-driven optimization mechanism. The paper proposes the first AI-augmented closed-loop quality engineering framework that integrates production feedback learning to enable cross-release adaptive evolution of quality strategies. The approach combines requirement feature mining, risk-driven test prioritization, defect prediction, and a constrained feedback mechanism informed by defect severity and production impact. Evaluated over six release cycles, the method consistently reduced defect leakage from 0.19 to 0.13, improved detection effectiveness from 0.72 to 0.84, and decreased test execution time by up to 35%, demonstrating stable and reproducible gains.
This study addresses the challenge of transforming stakeholder requirements into product requirements in software-driven automotive systems. Leveraging a dataset of 8,082 stakeholder requirements and 5,870 product requirements provided by Infineon, the research employs a hybrid methodology integrating structural statistics, decision modeling, traceability mining, textual analysis, and hardware-software linkage to systematically analyze the requirement refinement process. It reveals, for the first time, that requirement complexity primarily stems from ambiguous architectural scope and missing contextual information rather than linguistic redundancy. The work establishes a classification framework for mapping stakeholder to product requirements, identifies systematic differences across abstraction levels, and proposes key improvements in requirement validation, deviation management, and contextual tooling to support efficient and reusable automotive development.