Score
Designs, builds, and maintains software packages and the infrastructure that creates, stores, distributes, and installs them — including package formats, build recipes, metadata, versioning, signing, repository layout, and distribution pipelines. Implements and operates package management tools and publishing workflows (package managers, dependency resolution, repository servers, CI/CD for publishing) and handles platform- or language-specific packaging such as Linux distributions and Python (wheel/sdist and PyPI), while analyzing installation, upgrade, security, and compatibility issues.
Existing package managers suffer from semantic fragmentation due to language- and operating system-specific differences, making it difficult to precisely express cross-language dependencies, versioned system or hardware requirements, and hindering effective security vulnerability tracking. To address these challenges, this work proposes Package Calculus—the first unified formal model that captures the core mechanisms of mainstream package managers through semantic reduction. Serving as an intermediate representation, Package Calculus enables translation and resolution of dependencies across heterogeneous ecosystems. The model facilitates cross-language and cross-platform dependency interoperability and supports global analysis, thereby establishing a rigorous theoretical foundation and practical pathway for dependency resolution and security research.
This study addresses developer challenges arising from package manager lockfiles in reproducible builds, integrity verification, and routine maintenance. We conduct the first cross-ecosystem empirical investigation, systematically analyzing the design space of lockfiles across seven major package managers—including npm, pnpm, and Cargo—complemented by semi-structured interviews with 15 developers and documentary analysis of specifications. Our findings reveal critical discrepancies across ecosystems in semantic expressiveness, generation logic, and operational pain points. Based on these insights, we propose four developer-centric lockfile improvement principles—the first such human-centered guidelines in dependency management tooling—thereby bridging a gap in socio-technical research on dependency locking. The study provides an empirical foundation and actionable design guidance for next-generation lock mechanisms that are more maintainable, interpretable, and verifiable. (149 words)
Multilingual projects suffer from three core challenges: absence of cross-ecosystem dependency modeling, lack of versioning for external system/hardware dependencies, and poor interoperability among package managers. This paper introduces HyperRes—the first formal dependency resolution system that unifies multilingual and multisystem dependencies into a verifiable hypergraph model. Its contributions are threefold: (1) an environment-aware, versioned dependency model grounded in hypergraph theory, explicitly representing implicit system- and hardware-level dependencies; (2) a bidirectional metadata translation framework enabling zero-migration interoperability across dozens of package managers (e.g., npm, pip, apt); and (3) a hybrid solving strategy integrating constraint satisfaction problem (CSP) techniques with environment-specialized algorithms to achieve consistent, precise cross-ecosystem dependency resolution. Empirical evaluation demonstrates that HyperRes significantly improves reliability and reproducibility in multilingual environment construction.
The explosive growth of third-party packages in Linux distributions poses significant package management challenges. Method: This paper presents the first empirical study on the Package-to-Group (P2G) mechanism, leveraging large-scale data mining across 11,746 groups, 193,548 packages, and 89 distribution versions, complemented by cross-distribution comparative analysis and practitioner surveys. Contribution/Results: We propose a P2G research framework, define GValue—a quantitative metric for group quality—and identify six evolutionary patterns, including the counterintuitive finding that ungrouped packages are more prone to stagnation. We further categorize five package types highly likely to undergo P2G. Results show accelerating P2G adoption across mainstream distributions, yet widespread quality issues persist—particularly missing group descriptions and excessively small group sizes. Our findings yield actionable, community-oriented governance recommendations and practical guidelines for sustainable package grouping.
This paper presents a systematic review of core challenges in software deployment, including reproducibility, dependency resolution, trust mechanisms, and fine-grained incremental builds. It offers the first comprehensive evaluation of Nix’s pure functional approach across build systems, package management, system configuration, and development environments. Through dependency graph modeling and taxonomic analysis of related tools, the study clarifies Nix’s contributions to ensuring reproducible builds while exposing its limitations in trust establishment and incremental build support. The work further synthesizes cutting-edge community-driven solutions addressing these shortcomings and outlines promising directions for future research in reliable and efficient software deployment.
研究了跨生态系统软件包的普遍性、架构模式及其与项目健康度的关系,通过分析六大生态系统中的六百万个软件包,识别出五种架构模式。
To address poor reproducibility, low build efficiency, and insufficient deployment automation in embedded Linux system customization, this paper proposes a three-layer extensible architecture based on the Yocto Project. The architecture integrates GitLab CI and Docker to ensure environment isolation and enable continuous integration and deployment (CI/CD), while incorporating a local hash server (hashserv) and shared sstate cache server to significantly improve build artifact reuse. It supports automated real-time Linux kernel builds, QEMU-based simulation testing, and validation across six distinct boot scenarios. Experimental evaluation demonstrates substantial reduction in build time, markedly enhanced system stability and build reproducibility, and strong scalability and engineering deployability for industrial-grade applications.
This study addresses the sustainability of open-source Python libraries, which hinges critically on maintainer reachability, yet the availability and validity of associated email addresses have not been systematically assessed. We present the first large-scale empirical analysis of email contacts across 686,034 packages on PyPI and their corresponding GitHub repositories, integrating web scraping, email validation, and dependency graph construction to quantify the distribution and coverage of contact information. Our findings reveal that 81.6% of packages contain at least one valid email address, and 97.7% of transitive dependencies are reachable through valid contact information. Nevertheless, we identify over 698,000 invalid email records, highlighting both the current state of maintainer accessibility in the ecosystem and significant opportunities for improvement.
This study investigates the causal impact of SECURITY.md security policy files on the structure and evolution of software dependency chains within the PyPI ecosystem. Using a longitudinal dataset of 1,248 open-source projects, we construct temporal dependency trees and apply a quasi-experimental design combining difference-in-differences and propensity score matching to compare dependency management behaviors between projects with and without SECURITY.md. Results show that adopting SECURITY.md significantly enhances modularity: direct dependency count increases by 23.6%, and dependency update frequency rises by 31.4%, while transitive dependency depth remains unchanged. Late adopters exhibit stronger proactive dependency governance. This work provides the first empirical evidence that SECURITY.md is not merely a risk disclosure instrument but a key governance mechanism that strengthens software supply chain resilience—offering data-driven insights for evidence-based software supply chain security policy design.
This study addresses the widespread presence of package replicas in the Python Package Index (PyPI), which not only mislead developers but also serve as blind spots for known vulnerabilities and vectors for malware. Through a large-scale analysis of approximately 200,000 PyPI packages, the work integrates static code analysis, metadata comparison, and similarity detection, cross-referenced with vulnerability and malware databases to systematically uncover the prevalence and dual security risks of such replication practices. The research identifies 1,361 replica packages mimicking popular projects, 256 replicas harboring previously unknown vulnerabilities, and seven novel malicious replicas. Notably, it confirms that 4.79% of known malicious packages employ typosquatting or imitation of popular packages to carry out attacks.