Score
Designs, deploys, and operates Kubernetes clusters and platforms for running containerized workloads, covering cluster architecture and sizing, scheduling and resource policies, deployment manifests and operators, platform automation, CI/CD integration, and administration (including managed offerings such as Azure Kubernetes Service). Builds and debugs Docker container images, Kubernetes operators and internals, integrations and troubleshooting workflows, and analyzes performance, reliability, and operational issues for Kubernetes-based environments.
Adaptive performance management in cloud-native environments faces persistent challenges in jointly optimizing adaptability, elasticity, and efficiency. This paper systematically surveys 96 peer-reviewed publications from 2017 to 2023 and proposes a novel five-dimensional classification framework—spanning optimization objectives, control scope, decision-making mechanisms, automation levels, and validation methodologies. It is the first to holistically integrate reactive/predictive feedback loops, ML-driven resource forecasting, cross-dimensional benchmark datasets, and AIOps toolchains, identifying pattern-based adaptive architectures at the application layer as a critical research gap. Key findings include a marked surge in related work since 2023 and the consolidation of feedback control and machine learning as dominant paradigms. The study further releases a standardized validation dataset inventory—categorized by application, resource, and network dimensions—and a taxonomy of mainstream AIOps tools, thereby enabling reproducible, comparable experimental evaluation.
This study systematically reveals the significant impact of network misconfigurations on lateral movement attack risks in Kubernetes clusters. Addressing the limited coverage of existing detection tools, we propose a security assessment framework that integrates static configuration analysis with lateral movement path modeling. We conduct a large-scale, cross-organizational empirical study across 287 open-source applications, identifying— for the first time—634 real-world network misconfiguration vulnerabilities, far exceeding the detection capacity of mainstream tools. Our findings have driven remediation efforts in over 30 critical open-source projects; the proposed mitigation strategies have been adopted by multiple enterprises, substantially enhancing network isolation and overall security posture in production Kubernetes deployments.
In cloud-native systems, Kubernetes’ declarative configurations (YAML/Helm) impede architectural understanding, hindering developer and operator productivity. This paper introduces the first user-study-driven visualization framework that automatically and semantically faithfully maps raw Kubernetes resource specifications to interpretable architecture diagrams. Methodologically, it integrates a custom domain-specific language (DSL), the Kubernetes client API, and a graph-based encoding strategy to enable zero-intrusion integration and low-friction embedding into DevOps pipelines. Evaluated on three real-world systems, the tool accelerates architectural comprehension by 42% on average and improves modeling accuracy—reducing modeling errors by 68%. It has been adopted by the CNCF ecosystem as a recommended visualization tool. The core contributions are: (1) the first visualization generation paradigm explicitly designed for Kubernetes semantics; and (2) a solution that jointly ensures precision, scalability, and engineering practicality.
该研究比较了Docker容器与虚拟机在架构、性能、配置和安全方面的差异,分析了两者在隔离性与效率上的权衡,并提出混合架构作为解决方案。
This study addresses prolonged task completion times, low resource utilization, and high resource release latency in Docker/Kubernetes containers on cloud-native platforms running compute-intensive workloads (e.g., big data and deep learning). We systematically evaluate the performance impact of diverse resource scheduling strategies through system-level monitoring—leveraging cgroups and metrics-server—and multi-workload stress testing. For the first time, we empirically quantify how key resource configurations significantly affect task completion time (±79.4% variation) and resource release latency (+116.7% degradation). Based on these findings, we propose an evidence-driven configuration optimization paradigm that reduces maximum task completion time by up to 79.4% and precisely identifies configuration bottlenecks responsible for latency. Our results provide reproducible, transferable empirical foundations for resource management tuning and deployment decisions in cloud-native environments.
Research on containerization in multi-cloud environments remains fragmented, lacking a systematic, up-to-date synthesis. Method: We conduct a Systematic Mapping Study (SMS) spanning 2013–2024, analyzing 121 high-quality publications through bibliometric analysis, thematic coding, and ISO/IEC 25010 quality attribute modeling. Contribution/Results: We propose the first four-level classification framework—“Theme–Strategy–Quality Attribute–Tactic”—identifying four core research themes, 98 implementation strategies, 10 critical quality attributes, and 47 corresponding architectural tactics. Innovatively, we introduce a two-dimensional challenge-solution taxonomy organized along Security, Automation, Deployment, and Monitoring dimensions. This yields the first structured, reusable landscape of multi-cloud containerization, bridging theoretical research and industrial practice by supporting architecture design and technology selection—thereby addressing a longstanding gap in systematic knowledge integration for this domain.
This study addresses the problem of Kubernetes infrastructure drift, where runtime states deviate from architectural intent, by proposing an editable living architecture model. This model explicitly maps runtime facts to architectural designs, supports bidirectional synchronization between textual and graphical views, and enables non-intrusive Architecture-as-Code management through periodic consistency checks. The proposed approach is implemented using a subset of the Archer KDL, a VS Code-based prototype tool, and snapshot restoration techniques. Experimental evaluations conducted on three representative applications validate the feasibility of the method, demonstrating strong performance in both snapshot restoration accuracy and inconsistency detection recall.
This work addresses the limitations of Kubernetes’ default scheduler, which often leads to resource fragmentation and suboptimal utilization due to its local decision-making nature, while existing global scheduling approaches struggle with practical deployment in production clusters. The paper proposes OPSche, the first open-source plugin that collaboratively operates alongside the default scheduler by leveraging the Kubernetes scheduling framework to seamlessly integrate globally optimized schedules generated by external solvers through atomic validation and coordination hooks. OPSche supports three trigger modes—scheduling failure, periodic invocation, and queue stabilization—along with their blocking variants, thereby balancing scheduling quality, latency, and interference without replacing the native scheduler. Experimental results demonstrate that OPSche improves resource utilization by up to 3.0% across diverse cluster configurations and reduces scheduling latency by over one second.
This work addresses the limitations of traditional single-tenant batch systems in meeting the demands of AI training, secure computation on sensitive data, and mixed workloads requiring flexibility and reproducibility. It presents the first deployment of a multi-tenant Kubernetes infrastructure on the HPE Cray EX supercomputer (Isambard-AI), integrating trusted research environments with distributed AI model hosting services. The proposed architecture leverages KubeRay, Ray, vLLM, and HPE Slingshot interconnects to deliver a sandboxed, persistent platform. By extending Kubernetes beyond conventional cloud environments into high-performance bare-metal systems, this study demonstrates the feasibility of scalable “Kubernetes-as-a-Service” for national-scale AI infrastructure. It further identifies key implementation challenges and outlines an evolutionary pathway, offering a co-design paradigm for multi-tenant confidential computing in domains such as healthcare.
研究解决了Kubernetes安全配置错误问题,通过分析2,662个Stack Overflow问题构建分类体系,并使用大语言模型结合模式引导框架Kubecurity进行自动修复。
This study addresses the significant inconsistencies among current Kubernetes security hardening guidelines and configuration scanning tools in terms of recommendation coverage and risk scoring, which hinder effective security configuration decisions. The authors systematically analyze eight widely adopted hardening guides to derive a unified benchmark of 79 configuration recommendations and conduct a structured empirical evaluation of ten static scanning tools. For the first time, they establish a standardized framework for assessing Kubernetes configuration security. Their findings reveal substantial discrepancies in both coverage and risk assessment across existing guidelines and tools, underscoring the urgent need for a transparent and consistent security evaluation methodology. This work provides the community with a reproducible benchmark and actionable criteria for tool selection and policy alignment.