Score
Ability to design, implement, and maintain software components, libraries, or scripts using Java, C++, Python, or closely related languages; this includes writing idiomatic code, implementing algorithms and data structures, and integrating with language-specific ecosystems (build systems, package managers, and standard libraries). Involves unit testing and debugging, profiling and performance optimization, and producing readable, documented, and version-controlled code suitable for integration into larger systems.
研究了跨生态系统软件包的普遍性、架构模式及其与项目健康度的关系,通过分析六大生态系统中的六百万个软件包,识别出五种架构模式。
To address the challenges of simultaneously generating semantically consistent yet stylistically diverse multi-artifact programming exercises—namely source code, test specifications, and natural language descriptions—this paper proposes a compositional generation framework grounded in abstract syntax building blocks. The framework defines reusable syntactic abstractions and integrates templated mapping with multi-objective instantiation to ensure intent preservation and cross-modal co-generation. Its key innovations include: (i) enabling style-controllable, diverse outputs while guaranteeing semantic consistency; and (ii) providing a highly configurable generation interface that substantially reduces customization effort for new tasks. Experimental evaluation demonstrates that the approach outperforms existing baselines across three critical dimensions: generation quality, output diversity, and system extensibility.
Prior research on library migration has largely overlooked C/C++, leaving a critical gap in understanding its ecosystem’s evolution. Method: We construct the first large-scale, multi-source C/C++ library migration dataset, encompassing 19,943 projects across seven package managers, and conduct an empirical analysis integrating dependency graphs, commit histories, and issue trackers to systematically characterize migration behaviors, domains, and motivations—comparing findings against Python, JavaScript, and Java. Contribution/Results: We find C/C++ migrations concentrate in GUI, build-system, and OS development; 83.46% of source libraries map deterministically to a single target library; and unique drivers include reducing compilation time and unifying dependency management. Crucially, C/C++ migration patterns diverge significantly from dynamic languages (e.g., JS/Python) and Java—especially in domain distribution—demonstrating the dataset’s utility for developing specialized migration recommendation tools.
研究通过在JavaScript和Lua中实现Processing/p5,提出了一套软件决策指导原则,以解决不同编程语言环境下创意编码的一致性问题。
The absence of a standardized, sustainability-focused defect knowledge base for green software development hinders the advancement of automated sustainability analysis tools. Method: We propose the first systematic classification framework for sustainability weaknesses, derived through empirical analysis and pattern mining across ecological dimensions—including energy efficiency and resource waste—to semantically re-annotate and attribute code defects. Our approach explicitly decouples sustainability weaknesses from conventional software defect taxonomies (e.g., CWE), rigorously validating their non-transferability. Contribution/Results: We introduce the first standalone, ecology-aware sustainability weakness taxonomy, supported by formal modeling and empirical validation. The resulting knowledge base enables scalable, reusable foundations for static sustainability analysis, eco-conscious code optimization, and actionable sustainability recommendations in green software engineering.
This work addresses the challenges of IDE development posed by the rapid evolution of smart contract languages such as Move by presenting a high-performance IDE support system built atop the Move compiler and adhering to the Language Server Protocol (LSP). Through deep integration with existing language toolchains and the application of incremental parsing and optimized semantic analysis techniques, the system efficiently delivers rich IDE features even as the language undergoes continuous iteration. Deployed successfully within the Sui platform’s Move ecosystem, it significantly enhances developer experience and yields a reusable, evolution-aware IDE construction strategy applicable to other emerging programming language ecosystems.
This work addresses the challenge of balancing software quality, testability, and maintainability under rapid iteration and frequent requirement changes. It proposes Algorithm-Driven Development (ADD), a novel approach that unifies requirements specification and technical design by using algorithm flowcharts as a single, coherent artifact. This integration enables end-to-end modeling of requirements, architecture, and testing. Leveraging this model, the system automatically generates high-coverage acceptance tests and incorporates continuous integration with code coverage feedback. Industrial adoption at Dassault Systèmes demonstrates that ADD achieves over 95% code coverage, substantially reduces defect density, and ensures a stable delivery cadence, outperforming conventional test-driven development and test-after approaches.
This work addresses the significant performance degradation of large language models (LLMs) in generating code for constraint-based domain-specific languages (DSLs), such as OCL and Alloy, and the absence of systematic evaluation methodologies. The paper introduces the first evaluation framework tailored for constraint DSL code generation, which systematically assesses LLM capabilities in translating natural language to DSL through both syntactic correctness and semantic accuracy, leveraging formal verification. Experimental comparisons across Python, OCL, and Alloy reveal that LLMs perform markedly better on general-purpose languages, that models with limited context windows struggle to jointly generate constraints and domain models, and that incorporating code repair and multi-candidate generation strategies substantially improves output quality. The framework further enables systematic analysis of prompting templates, repair mechanisms, and multi-turn generation strategies.
This work addresses the frequent neglect of sampling strategy design and generalizability in software engineering research, which often undermines the representativeness of empirical findings. To remedy this, the paper introduces a domain-specific language (DSL) that explicitly models complex sampling workflows over code repositories through composable sampling operators, enabling—for the first time—formal specification and reasoning about the generalizability of sampling strategies. Implemented as a fluent Python API, the DSL is integrated with a statistical metric system to quantitatively assess the external validity of sampled datasets. The authors demonstrate the expressiveness and practical utility of their approach by reconstructing and formalizing the sampling procedures from multiple Mining Software Repositories (MSR) studies, thereby validating the framework’s capacity to capture real-world methodological diversity.
This study addresses the emerging yet underexplored sustainability implications of code generated by large language models (LLMs), which—when deployed at scale—can incur significant energy consumption and environmental impact due to inefficiencies. Through a systematic literature review, this work provides the first structured synthesis of existing research, critically examining key dimensions such as prompt engineering, fine-tuning strategies, and energy-efficiency evaluation metrics. The analysis reveals a critical lack of consensus on defining sustainability in this context, alongside the absence of standardized measurement methodologies and benchmarking frameworks. The paper calls for establishing a coherent research paradigm and a unified evaluation framework to guide the development of future LLM-based code generation systems toward greater efficiency and environmental responsibility.