Score
Designs and implements reusable, readable, and configurable Python code—ranging from small scripts and libraries to end-to-end user tools—that implement algorithms, data-processing pipelines, and visualizations. Builds validation and reproducibility into those tools through testing, example case studies, and configurable interfaces for deployment and use.
Poor code readability in scientific software severely hinders cross-team collaboration and research reproducibility—particularly among self-taught researchers, who typically lack formal training in readability best practices, resulting in opaque naming conventions and inadequate documentation. This study employs a mixed-methods approach—including surveys, in-depth interviews, and statistical analysis—across 57 interdisciplinary researchers to empirically investigate current practices. It reveals, for the first time, that in the absence of structured training, researchers heavily rely on informal, ad hoc commenting practices; further, it identifies large language models (LLMs) as an emerging paradigm for enhancing code quality. Results show that 57.9% of participants received no readability-specific instruction, with inconsistent naming and missing documentation identified as the two primary bottlenecks. Based on these findings, we propose a lightweight, human-centered code quality support framework tailored for scientific programmers—addressing a critical gap in the human factors literature on scientific code readability.
This study systematically evaluates large language models—particularly GPT-4—for idiomatic Python code refactoring to enhance clarity, efficiency, and readability. We propose a prompt-engineering–based automated refactoring recommendation method that jointly identifies non-idiomatic patterns and generates idiomatic alternatives. Evaluation combines human assessment with static analysis tools (e.g., Pylint, Vulture) to benchmark accuracy, coverage, and contextual adaptability. Our empirical evaluation demonstrates, for the first time, that GPT-4 significantly outperforms traditional static analysis baselines in both recommendation accuracy and scope—especially in semantically nuanced, context-dependent refactoring scenarios requiring deep program understanding. A randomized human validation sample yields a 92.3% correctness rate for GPT-4’s suggestions, confirming its viability as a high-precision, context-aware assistant for idiomatic refactoring. The work establishes LLMs as robust, adaptive tools for practical, semantics-driven code improvement.
AI-generated code exhibits low reliability, high maintenance overhead, and poor verifiability in embedded systems. To address these challenges, this paper introduces Pythoness—a domain-specific language (DSL) designed for embedded development that enables developers to specify behavioral requirements via natural-language descriptions or formal tests, thereby replacing low-level coding and enabling collaborative programming with large language models (LLMs). Its core contribution is the first proposal of a test-driven LLM programming paradigm: unit and property tests guide prompt engineering, establishing a closed-loop workflow of code generation, execution, and feedback, while runtime verification continuously ensures correctness. Evaluation of a prototype implementation demonstrates that Pythoness significantly improves test pass rates (+32%) and reduces defect density (−41%) compared to specification-only approaches, while also enhancing maintainability and formal verifiability.
This work addresses the risk that automated Python refactoring tools may inadvertently introduce behavioral changes, thereby compromising software reliability. To tackle this issue, the authors propose a novel approach that leverages foundation models as semantic oracles, integrated with Git diff parsing and automated validation, to detect behavior-altering refactorings. Applying this method to 217 refactoring instances produced by the Rope tool, the study uncovers 13 previously unknown defects, 12 of which have been acknowledged and fixed by the developers. This demonstrates the effectiveness of the technique in enhancing the trustworthiness and practical utility of automated refactoring tools.
This paper addresses persistent software engineering (SE) challenges in Jupyter Notebooks—including low code reusability, poor readability, unreliable execution environments, and weak long-term accessibility—through a systematic literature review (SLR) of 146 studies published through December 2024. The analysis reveals that human-computer interaction (HCI) researchers dominate publication, with only 64 studies providing reusable links—and most notebooks absent from permanent repositories. Core SE concerns such as testing, refactoring, and documentation lack notebook-specific solutions. This work constitutes the first comprehensive identification of notebook-native SE challenges and proposes three novel research directions: (1) automated cell-level unit testing, (2) cross-notebook refactoring and clone detection, and (3) cell-granularity collaborative documentation generation. The findings establish an empirical foundation and technical roadmap for developing notebook-native SE methodologies.
本文通过大规模实证研究,分析了Python项目跨操作系统的移植性问题,并提出分类方法和修复模式,以提高开发者的应对能力。
This study addresses the challenges of aligning Python programs with formal specifications and verifying backward compatibility following component updates. To this end, this work proposes a source-preserving verification framework that expresses contracts through Python annotations and automatically generates proof artifacts in Dafny and Lean, thereby unifying development and functional verification at the source-code level. Furthermore, relational product techniques are introduced to check compatibility after component updates. The primary contribution is an end-to-end workflow bridging source-level proofs and compatibility verification, ensuring rigorous formal guarantees while preserving executable semantics. The associated codebase has been made publicly available.
研究针对Python应用中的本地代码bug问题,通过分析216个真实项目中的案例,揭示了这些bug的症状、原因及修复策略。
Automatically reproducing executable bug-fix code pairs from unstructured developer Q&A posts is hindered by ambiguous descriptions and missing dependencies. This work proposes Reprodgen, the first end-to-end automated framework that leverages large language models to jointly model code intent (CI), functional requirements (FR), and structured chains of thought (SCoT) to generate semantically consistent and executable bug-fix code pairs. The approach incorporates an LLM-based iterative review mechanism coupled with real execution validation to ensure correctness. Evaluated on Stack Overflow and GitHub Issues across seven widely used data science libraries, the study introduces the first expert-validated, runnable benchmark of bug-fix pairs. Experimental results demonstrate that Reprodgen reliably reproduces code pairs exhibiting clear behavioral differences between buggy and fixed versions.
研究分析了1000个GitHub仓库中Python库和框架对类型提示的采用、维护情况,通过提取类型注解等方法探讨其使用模式及演变。