Score
Designs, builds, and configures container images and containerized application environments that package software and its dependencies for isolated, reproducible execution. Implements and analyzes container runtime behavior, orchestration, networking, storage, resource limits, lifecycle (start, update, scale, stop), and security settings to enable deployment and maintenance.
Containerization enhances operational efficiency but intensifies multidimensional security challenges—including runtime protection, network isolation, configuration compliance, software supply chain security, and monitoring-response capabilities. To address these, this paper proposes a five-dimensional collaborative governance model for production-grade container security, deeply integrating DevSecOps across the entire lifecycle and transcending traditional perimeter-based defense paradigms. Methodologically, the model unifies eBPF-based real-time runtime detection, OCI image signature verification, SBOM-driven supply chain auditing, zero-trust network policy enforcement, and a tightly coupled Prometheus–Falco incident response mechanism. Evaluated on mainstream cloud-native platforms, the approach reduces critical misconfigurations by 92%, shortens mean vulnerability response time to 3.7 minutes, and enables construction of a CNCF Sig-Security-certified hardened baseline—delivering a practical, layered defense framework for containerized environments.
This study addresses the practice gap in security management for containerized software development. Through two rounds of semi-structured interviews with 35 practitioners, it systematically investigates frontline engineers’ perceptions of security risks, mitigation strategies, and implementation barriers in Docker and Kubernetes environments. Applying thematic coding and cross-case comparison, the study uniquely integrates technical and non-technical dimensions to identify 12 high-frequency security challenges. It further distills seven technical enablers—such as Software Bill of Materials (SBOM) integration and runtime policy engines—and five non-technical enablers—including cross-functional collaboration mechanisms. By bridging the “practice-awareness–technical-implementation” divide in container security, this work fills a critical gap in software engineering research. It provides empirically grounded foundations and design principles for developing practical, adoptable container security governance frameworks.
In software development, manual or scripted environment configuration is inefficient and unreliable—especially when onboarding unfamiliar Python codebases. This paper introduces Repo2Run, the first end-to-end LLM agent that fully automates environment setup: from source-code analysis and dependency inference to generating executable Dockerfiles. Its key contributions are: (1) an atomic configuration synthesis mechanism, leveraging dual-environment isolation and rollback support to guarantee execution atomicity; and (2) a structured Dockerfile generator guided by LLM-based planning and iterative sandbox execution feedback, enabling precise translation of successful configuration steps into robust, reproducible image definitions. Evaluated on a benchmark of 420 real-world Python repositories, Repo2Run achieves an 86.0% configuration success rate—outperforming the best baseline by 63.9 percentage points.
Although Docker is widely assumed to ensure reproducibility of software environments, its practical efficacy remains insufficiently validated. This study presents the first systematic investigation combining a literature review with large-scale empirical analysis of 5,298 real-world GitHub projects. By reconstructing Docker images, performing differential comparisons, and mining workflow patterns, we quantitatively assess the reproducibility of Docker builds and the effectiveness of recommended best practices. Our findings reveal that a significant proportion of Docker builds are not reproducible, and existing best practices offer limited improvements in practice. These results challenge the prevailing assumption that “containers guarantee reproducibility” and provide empirical evidence and actionable insights for enhancing reproducibility in computational research.
Research on containerization in multi-cloud environments remains fragmented, lacking a systematic, up-to-date synthesis. Method: We conduct a Systematic Mapping Study (SMS) spanning 2013–2024, analyzing 121 high-quality publications through bibliometric analysis, thematic coding, and ISO/IEC 25010 quality attribute modeling. Contribution/Results: We propose the first four-level classification framework—“Theme–Strategy–Quality Attribute–Tactic”—identifying four core research themes, 98 implementation strategies, 10 critical quality attributes, and 47 corresponding architectural tactics. Innovatively, we introduce a two-dimensional challenge-solution taxonomy organized along Security, Automation, Deployment, and Monitoring dimensions. This yields the first structured, reusable landscape of multi-cloud containerization, bridging theoretical research and industrial practice by supporting architecture design and technology selection—thereby addressing a longstanding gap in systematic knowledge integration for this domain.
Container technologies are widely adopted, yet their full lifecycle entails significant security risks; existing software engineering literature lacks systematic, empirically grounded integration of container security knowledge. To address this gap, we conducted a systematic mapping study (SMS) complemented by bibliometric analysis and thematic coding across 129 empirical studies. Our work introduces the first structured, evidence-based taxonomy of security risks for containerized systems—identifying 23 core risk categories and vulnerabilities, explicating their root causes and impacts, and synthesizing reusable mitigation strategies. Additionally, we catalog 47 security practices and tools. The taxonomy enables cross-phase risk mapping—from development through deployment—and integrates fragmented knowledge into a coherent framework. It establishes a theoretical benchmark for container security research and delivers actionable, engineering-oriented guidance for practitioners.
该研究比较了Docker容器与虚拟机在架构、性能、配置和安全方面的差异,分析了两者在隔离性与效率上的权衡,并提出混合架构作为解决方案。
While large language models (LLMs) can generate executable multi-service application environments, they often deviate from the architectural and security requirements essential for production deployment. This work proposes a method to automatically generate Dockerfiles and Docker Compose configurations solely from code repositories, evaluating deployment fidelity through end-to-end HTTP testing and structural comparison. It explicitly distinguishes between functional correctness and fidelity to deployment intent, deriving a minimal set of explicit deployment specifications that cannot be inferred automatically from source code alone. Experiments successfully reproduce the topology and dependencies of three heterogeneous multi-service systems, confirming functional feasibility; however, critical production-grade features—such as network isolation and multi-stage builds—are consistently absent, revealing fundamental limitations in current LLMs’ ability to model deployment intent.
Modern software systems face a codebase organization dilemma: monorepos ensure consistency but suffer from poor scalability and complex toolchains, whereas multi-repos improve modularity at the cost of increased dependency coordination and integration overhead. To address this, we propose Causify Dev—a novel development system introducing the “Runnable Directory” paradigm, wherein each directory functions as an isolated, self-contained execution unit with its own dependencies and full lifecycle management. Leveraging lightweight unified development environments and containerized workflows—built on Docker, standardized CI/CD pipelines, and shared utility libraries—the system achieves end-to-end decoupled collaboration while preserving monorepo-level consistency and multi-repo-style modularity. Empirical evaluation demonstrates that Causify Dev significantly enhances reliability, maintainability, and engineering scalability for large-scale codebases, while simultaneously reducing toolchain complexity and cross-team coordination costs.
This study addresses the limited understanding of containerization practices in machine learning (ML) projects, particularly the lack of systematic investigation into how iterative ML workflows affect Docker image size, build performance, and caching behavior. Through a large-scale empirical analysis of Dockerfiles from 1,993 open-source ML projects—integrating static parsing, build log tracing, cache behavior monitoring, and semantic mining of code commits—the work reveals ML-specific container usage patterns and proposes seven ML-tailored Dockerfile refactoring strategies. The findings show that ML images average 10.27 GB in size and require 8.84 minutes to build; 44.4% of commits trigger rebuilds, with 96.4% caused by context changes and 71% exhibiting computational redundancy. The proposed methods substantially reduce image size and improve build efficiency.