Score
Designs, implements, and maintains production-quality software using Python, including scripts, command‑line tools, libraries, backend services, system-level programs, and data processing components. Builds and hardens maintainable, testable, and deployable Python code by applying packaging, dependency management, performance tuning, observability, and production deployment practices to create reliable, scalable, and reusable systems and pipelines.
本文通过大规模实证研究,分析了Python项目跨操作系统的移植性问题,并提出分类方法和修复模式,以提高开发者的应对能力。
研究通过ImportMine分析Python导入相关漏洞和错误,结合安全公告与PyPI历史数据,识别并分类问题,提出修复方案。
This study addresses the challenges of aligning Python programs with formal specifications and verifying backward compatibility following component updates. To this end, this work proposes a source-preserving verification framework that expresses contracts through Python annotations and automatically generates proof artifacts in Dafny and Lean, thereby unifying development and functional verification at the source-code level. Furthermore, relational product techniques are introduced to check compatibility after component updates. The primary contribution is an end-to-end workflow bridging source-level proofs and compatibility verification, ensuring rigorous formal guarantees while preserving executable semantics. The associated codebase has been made publicly available.
This work addresses the risk that automated Python refactoring tools may inadvertently introduce behavioral changes, thereby compromising software reliability. To tackle this issue, the authors propose a novel approach that leverages foundation models as semantic oracles, integrated with Git diff parsing and automated validation, to detect behavior-altering refactorings. Applying this method to 217 refactoring instances produced by the Rope tool, the study uncovers 13 previously unknown defects, 12 of which have been acknowledged and fixed by the developers. This demonstrates the effectiveness of the technique in enhancing the trustworthiness and practical utility of automated refactoring tools.
Addressing the challenges of implementing FAIR principles and open science practices in research-oriented Python software engineering—particularly low automation and poor adoption—the paper introduces the first end-to-end automated framework tailored for scientific computing. Methodologically, it integrates Configuration-as-Code, containerized DevOps, and domain-specific academic software engineering best practices within a cloud-native GitHub ecosystem. The framework enables one-click generation of PEP 517/518-compliant buildable package scaffolds, Sphinx-based documentation, pytest test suites, and GitHub Actions CI/CD pipelines. Leveraging Cookiecutter templates and a RESTful API control center, it ensures zero-friction onboarding for both new and legacy projects. Empirical evaluation demonstrates a 90% reduction in project initialization time, achieves 100% baseline test coverage and documentation completeness, and has been validated across multiple prominent open-source scientific libraries; its template repository is widely adopted by the community.
This study addresses the lack of systematic empirical research on the quality and security implications of code refactoring submitted by AI agents in real-world software projects. It presents the first quantitative analysis of AI-generated Python refactoring pull requests from the AIDev dataset, evaluating changes across maintainability, code quality, and security dimensions using tools such as PyQu, Pylint, and Bandit. The work further establishes a mapping between 24 common refactoring operations and potential issues. Findings reveal that 22.5% of changes improved code quality—primarily usability—yet 24.17% introduced new Pylint violations and 4.7% introduced security vulnerabilities. Although 73.5% of pull requests were merged, indicating high developer acceptance, the results underscore an urgent need for stronger quality and security gating mechanisms in AI-assisted development workflows.
研究针对Python应用中的本地代码bug问题,通过分析216个真实项目中的案例,揭示了这些bug的症状、原因及修复策略。
This study addresses the sustainability of open-source Python libraries, which hinges critically on maintainer reachability, yet the availability and validity of associated email addresses have not been systematically assessed. We present the first large-scale empirical analysis of email contacts across 686,034 packages on PyPI and their corresponding GitHub repositories, integrating web scraping, email validation, and dependency graph construction to quantify the distribution and coverage of contact information. Our findings reveal that 81.6% of packages contain at least one valid email address, and 97.7% of transitive dependencies are reachable through valid contact information. Nevertheless, we identify over 698,000 invalid email records, highlighting both the current state of maintainer accessibility in the ecosystem and significant opportunities for improvement.
This study addresses version conflicts, interpreter incompatibilities, and inefficient backtracking in Python dependency resolution by constructing a PyPI dependency knowledge graph and proposing an interpreter-aware SMT reasoning technique. By jointly encoding package dependencies and interpreter constraints into SMT formulas, this approach overcomes the limitations of traditional blind search methods, enabling precise co-resolution of dependencies and runtime environments. Experimental results demonstrate that the proposed method achieves speedups of 6.9× and 9.6× over pip and Conda, respectively. Furthermore, it consistently generates constraint-consistent executable environments, significantly enhancing both the efficiency and reliability of dependency resolution in complex Python ecosystems.
Automatically generating verifiable Python formal specifications remains challenging, and developers often abandon automated verification tools due to the tediousness of manually writing contracts. This work proposes a closed-loop approach that integrates large language models with symbolic execution (CrossHair) to automatically generate and iteratively refine icontract-style contract annotations without modifying the original code. The method leverages feedback from symbolic execution to drive specification refinement and simultaneously produces coverage-guided pytest stubs and debugging artifacts. Experimental results demonstrate that the approach successfully generates CrossHair-compatible specifications for most programs, significantly enhancing the practical feasibility of automated verification, while also revealing real-world limitations arising from the boundaries of symbolic exploration and behavioral discrepancies in large language models.