Score
Designs and produces reusable, parameterized code templates and example projects that developers can adapt for common tasks, including end-to-end templates and migration examples. Creates diagnostic and failure‑mode snippets, standardized patterns, and templating conventions to make code reuse, onboarding, and automated generation reliable and maintainable.
This study addresses the lack of systematic understanding regarding the application domains, maintenance characteristics, and effective design practices of GitHub template repositories. Conducting the first large-scale empirical investigation, the work integrates data mining, statistical analysis, code quality assessment tools—detecting code smells, vulnerabilities, and security hotspots—and an LLM-as-a-judge classification approach to systematically uncover domain distributions, language-specific quality variations, and maintenance patterns. The findings reveal web development as the dominant application domain, with high-quality templates consistently adhering to software engineering best practices and offering comprehensive documentation. Through qualitative evaluation, the study distills actionable design guidelines and identifies common pitfalls, providing practical guidance for developers creating or using template repositories.
This work addresses the common neglect of software design principles in existing automated code generation approaches, which often results in mobile applications with poor architectural quality. To overcome this limitation, the authors propose a novel method that integrates software product line engineering with variability modeling of design patterns. For the first time, the Universal Variability Language (UVL) is employed to explicitly capture structural and behavioral variations of design patterns, enabling their integration as configurable assets within the code generation pipeline. Leveraging UVL models, the Jinja templating engine, and Swift-based code synthesis, the proposed system supports the customizable, automated generation of design patterns such as Singleton and Strategy. This approach not only preserves architectural integrity but also significantly enhances application maintainability and reusability.
Template engine applications are notoriously difficult to debug and repair due to characteristics such as mixed-language composition, opaque data flows, and delayed validation, yet research in this area has long been scarce. This work presents the first large-scale empirical study of 1,004 real-world defects across 15 widely used template engines, systematically characterizing typical symptoms—predominantly abnormal rendering (48.61%)—identifying 17 root cause categories, and revealing collaborative repair patterns spanning both templates and host code (67.92% of fixes confined to templates, while over 20% require modifications to host logic). Based on these findings, the study offers actionable recommendations for developers and tool designers and implements two prototype debugging tools for the Jinja engine, demonstrably enhancing development and debugging efficiency for template-based applications.
This study addresses the challenges of high cost, error-proneness, and defect propagation in cross-repository code and test reuse during software refactoring. Through action research, the authors conduct bidirectional empirical analyses on real-world cases such as Soot/SootUp and FindBugs/SpotBugs, identifying for the first time the bidirectional reuse requirements and semantic reuse patterns inherent in refactoring scenarios. They propose a semantic alignment–based code mapping approach coupled with a hierarchical, extensible clone detection mechanism. Experimental results demonstrate that their method reduces irrelevant clones by 33%–99% on average and achieves a benchmark precision of 86%. The practical impact is further evidenced by five reported issues and ten pull requests submitted to open-source communities, eight of which have already been merged, confirming the approach’s effectiveness and applicability.
In software design, paradigm-implied semantic expectations—such as data abstraction consistency and feedback-control closed-loop behavior—are often left implicit, leading to design deviations and verification challenges. To address this, we introduce the concept of *design obligations*: explicit, logically formalizable, and verifiable specifications that codify such implicit constraints inherent to design paradigms. Leveraging formal modeling and paradigm semantics analysis, we establish two obligation frameworks—one for data-abstraction-based systems and another for feedback-driven adaptive systems—precisely capturing their core semantic requirements. We demonstrate that common design flaws stem from obligation violations and show how these obligations enable rigorous compliance verification and pedagogical application. This work bridges the semantic gap between design intent and implementation, providing both theoretical foundations and a methodological framework for paradigm-driven design assurance.
This work addresses the lack of a scalable, traceable, and systematic approach to modernizing large-scale legacy systems while preserving both functional and non-functional characteristics. The authors propose a four-phase model-driven method that leverages a semantically rich intermediate model to uniformly abstract a legacy system’s structure, dependencies, and metadata. By designing semantics-preserving transformation rules, the approach enables semi-automated migration to modern platforms such as web-based architectures. The method establishes an end-to-end model-driven pipeline that integrates semantic metadata modeling with automated code synthesis. Evaluated on an industrial-scale .NET system, it successfully migrated core UI components, significantly enhancing maintainability and scalability while reducing modernization risks and manual effort.
This work addresses the fragility of traditional requirements traceability in safety-critical software development, where reliance on external documentation often leads to silent breakdowns as code, requirements, and tests evolve independently. To overcome this, the paper proposes internalizing traceability as an intrinsic property of code structure by introducing language-native “Traceable” elements that enable compile-time verification of bidirectional links among requirements, implementations, and tests. The approach integrates code generation, metadata embedding, and build-time validation into the development workflow, providing proactive traceability assurance. When requirement changes disrupt traceability chains, the system automatically triggers warnings or build failures, thereby preventing traceability decay and significantly enhancing maintainability and reliability throughout software evolution.
Existing automated test generation approaches heavily rely on code coverage metrics, often failing to capture diverse test scenarios driven by implicit requirements and thereby risking the omission of critical defects. To address this limitation, this work proposes TestGeneralizer, a novel framework that treats an initial test as an executable specification. By leveraging large language models to infer its underlying requirements, TestGeneralizer constructs reusable test scenario templates and systematically instantiates them, enabling generalization from a single test to a comprehensive set of scenarios. This paradigm shifts beyond traditional coverage-driven methods, and empirical evaluation on twelve open-source Java projects demonstrates that TestGeneralizer improves mutation detection rate by 31.66% and LLM-assessed scenario coverage by 23.08% compared to ChatTester.
This study addresses the challenge of automatically detecting software design patterns in source code to support architectural understanding and quality assessment. It presents the first systematic evaluation of four large language models—including NextCoder and Gemma 3—as well as two ensemble strategies combining three models, for recognizing five classic design patterns: Singleton, Adapter, Bridge, Composite, and Decorator. The work investigates the impact of three input modalities—raw source code, PlantUML diagrams, and textual descriptions—on detection performance. Experimental results demonstrate that NextCoder and Gemma 3 achieve the highest accuracy among individual models, while ensemble approaches further enhance performance, thereby confirming the effectiveness and potential of large language models in design pattern recognition tasks.