Score
Designs and partitions systems into self-contained, non-cooperating modules with strict interfaces, assigning local processing responsibilities to each module and structuring data flows to minimize inter-module communication. Builds and evaluates modular partitioning schemes and implementations on existing hardware stacks to enforce hard boundaries and reduce cross-module dependencies.
Microservices achieve physical isolation but fail to prevent the proliferation of logical coupling, undermining module independence. This paper proposes a novel modularization paradigm based on universal interface boundaries, constructs a quantifiable model for assessing module independence, and designs a runtime mechanism supporting dynamic loading, unloading, and hot updates within a single process. Its core contributions are: (1) reframing module independence as a formal, modelable, and measurable system property—moving beyond qualitative assertions; (2) replacing implicit dependencies with explicit interface contracts to fundamentally block coupling propagation; and (3) implementing the EIGHT platform prototype, which achieves microservice-level module autonomy within a monolithic process. Experimental results demonstrate that the approach significantly reduces the impact scope of cross-module changes, enhancing system maintainability and evolutionary efficiency. It provides both theoretical foundations and practical pathways for next-generation architectures transcending the monolith–microservice dichotomy.
Heterogeneous hardware environments exhibit uneven resource distribution, rendering conventional lock mechanisms performance bottlenecks; existing solutions typically target single hardware types and fail to coordinate heterogeneous resources effectively. This paper introduces Modular Lock Decomposition—a novel paradigm that decouples lock functionality into independent, deployable modules (e.g., acquisition, waiting, wakeup) and dynamically assigns each module to appropriate hardware components (e.g., CPU cores, GPUs, FPGAs, cache levels) based on their architectural characteristics, enabling fine-grained, cross-architecture resource adaptation. To our knowledge, this is the first systematic shift in lock design from monolithic structures to hardware-aware modular architectures. Experimental evaluation under typical concurrent workloads demonstrates an average 42% reduction in lock contention latency and a 1.8× throughput improvement, significantly enhancing lock scalability and heterogeneous resource utilization.
This study addresses the NP-hard problem of hardware/software partitioning in computing architectures. Leveraging the directed pathwidth of task graphs, this work proposes a novel family of problem formulations that subsumes existing models, along with exact fixed-parameter tractable (FPT) algorithms. Methodologically, by integrating directed pathwidth analysis, FPT theory, and integer linear programming (ILP), the proposed approach achieves exact and efficient solutions for this problem family. The primary theoretical contribution lies in extending the modeling framework for hardware/software partitioning and establishing its fixed-parameter tractability. Empirically, experiments on real-world application scenarios demonstrate that the proposed method achieves up to a 200-fold speedup over general-purpose ILP solvers such as Gurobi, highlighting its practical efficacy and computational advantage.
本文通过引入Griotte和Griotte OS,利用形式化方法验证了CHERIoT基于能力的隔离机制的有效性和安全性。
Traditional operating systems suffer from poor scalability on many-core processors and low parallel efficiency due to their inability to perceive application semantics. To address this, we propose NetworkedOS—a novel application-aware, networked OS architecture. Our approach leverages compile-time dynamic instruction dependency analysis to construct a multi-layer network model that explicitly captures runtime dependencies among applications, the kernel, and hardware. We further design an overlapping graph partitioning algorithm to jointly optimize parallel execution and inter-core communication overhead, and implement a runtime process affinity mapping scheduler. Crucially, NetworkedOS breaks the conventional “black-box” OS assumption regarding application semantics for the first time. Experimental evaluation shows that NetworkedOS achieves a 7.11× speedup over Linux on a 128-core system and a 2.01× improvement over Barrelfish on a 64-core system, significantly enhancing scalability and resource utilization under large-scale parallel workloads.
This work addresses the poor scalability of existing memory isolation mechanisms in supporting large-scale, fine-grained intra-process security domains by proposing an efficient linked isolation model built upon CHERI hardware capabilities. We introduce a novel "one-click" library-boundary isolation strategy, combined with deep customizations to compilers and operating systems for both Armv8-A and RISC-V architectures. This approach enables fine-grained isolation of tens of thousands of security domains and over 500 concurrent domains within UNIX user space, substantially surpassing the concurrency limitations of mechanisms such as Intel MPK. The proposed solution has been validated on Morello and Codasip X730 platforms, demonstrating that the vast majority of C/C++ programs execute without modification, requiring only minor adaptations for the V8 engine.
Existing hypergraph partitioning methods often become trapped in local optima, limiting partition quality. This work proposes ComPart, a novel framework that integrates community structure guidance during both the initial partitioning and uncoarsening phases. It is the first to comprehensively incorporate community detection throughout the entire uncoarsening process and extends the theory of local dense decomposition from graphs to hypergraphs to generate high-quality initial partitions. By synergistically combining diverse community detection techniques, hypergraph local dense decomposition, and a multilevel partitioning strategy, ComPart consistently outperforms state-of-the-art methods on standard benchmarks, achieving significantly improved partition quality.
This study addresses the coordination challenges arising from independent control across cloud, high-performance computing (HPC), and edge AI infrastructures. Conceptualizing the AI platform as a "system of systems," this work proposes an architectural paradigm characterized by usage fusion and federated control. Methodologically, it adopts a systems engineering framework that achieves cross-domain coordination through interface contracts while preserving native control planes. The approach incorporates boundary testing, responsibility models, and seven integration facets, leveraging interface mapping, policy contexts, and operational evidence to guide integration design. The primary contribution lies in establishing a unified framework for evaluating interoperability, governance capabilities, and fault isolation, thereby delineating clear directions for future research.
This work addresses the programming challenges and error-proneness introduced by Arm’s POE2 architecture, which employs a complex spatiotemporal permission mechanism yet lacks a unified security model. We propose the first general-purpose secure programming model tailored for POE2, abstracting away the intricacies of its spatial and temporal indexing and encapsulating hardware features such as memory protection keys, dedicated registers, and table structures. By doing so, our model significantly simplifies permission management while preserving POE2’s strong security guarantees. It naturally supports common intra-process isolation patterns used in software partitioning, enabling developers to construct secure isolated systems more efficiently and with fewer errors.
This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.