design hard-boundary modules

Designs and partitions systems into self-contained, non-cooperating modules with strict interfaces, assigning local processing responsibilities to each module and structuring data flows to minimize inter-module communication. Builds and evaluates modular partitioning schemes and implementations on existing hardware stacks to enforce hard boundaries and reduce cross-module dependencies.

designhard-boundarymodules

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Microservices Are Dying, A New Method for Module Division Based on Universal Interfaces

Nov 06, 2025
QW
Qing Wang
🏛️ BNRist | Tsinghua University

Microservices achieve physical isolation but fail to prevent the proliferation of logical coupling, undermining module independence. This paper proposes a novel modularization paradigm based on universal interface boundaries, constructs a quantifiable model for assessing module independence, and designs a runtime mechanism supporting dynamic loading, unloading, and hot updates within a single process. Its core contributions are: (1) reframing module independence as a formal, modelable, and measurable system property—moving beyond qualitative assertions; (2) replacing implicit dependencies with explicit interface contracts to fundamentally block coupling propagation; and (3) implementing the EIGHT platform prototype, which achieves microservice-level module autonomy within a monolithic process. Experimental results demonstrate that the approach significantly reduces the impact scope of cross-module changes, enhancing system maintainability and evolutionary efficiency. It provides both theoretical foundations and practical pathways for next-generation architectures transcending the monolith–microservice dichotomy.

Addressing dependency propagation in microservices through module independence calculationDeveloping EIGHT platform architecture enabling dynamic runtime modifications in monolithic applicationsProposing universal interfaces to eliminate inter-module dependencies in system design

Towards Lock Modularization for Heterogeneous Environments

Aug 11, 2025
HZ
Hanze Zhang
🏛️ Shanghai Jiao Tong University

Heterogeneous hardware environments exhibit uneven resource distribution, rendering conventional lock mechanisms performance bottlenecks; existing solutions typically target single hardware types and fail to coordinate heterogeneous resources effectively. This paper introduces Modular Lock Decomposition—a novel paradigm that decouples lock functionality into independent, deployable modules (e.g., acquisition, waiting, wakeup) and dynamically assigns each module to appropriate hardware components (e.g., CPU cores, GPUs, FPGAs, cache levels) based on their architectural characteristics, enabling fine-grained, cross-architecture resource adaptation. To our knowledge, this is the first systematic shift in lock design from monolithic structures to hardware-aware modular architectures. Experimental evaluation under typical concurrent workloads demonstrates an average 42% reduction in lock contention latency and a 1.8× throughput improvement, significantly enhancing lock scalability and heterogeneous resource utilization.

Addressing lock inefficiency in heterogeneous hardware environmentsOvercoming resource bottlenecks in distributed lock operationsProposing modular locks for optimized hardware resource utilization

This study addresses the NP-hard problem of hardware/software partitioning in computing architectures. Leveraging the directed pathwidth of task graphs, this work proposes a novel family of problem formulations that subsumes existing models, along with exact fixed-parameter tractable (FPT) algorithms. Methodologically, by integrating directed pathwidth analysis, FPT theory, and integer linear programming (ILP), the proposed approach achieves exact and efficient solutions for this problem family. The primary theoretical contribution lies in extending the modeling framework for hardware/software partitioning and establishing its fixed-parameter tractability. Empirically, experiments on real-world application scenarios demonstrate that the proposed method achieves up to a 200-fold speedup over general-purpose ILP solvers such as Gurobi, highlighting its practical efficacy and computational advantage.

computational cost optimizationHardware-Software Partitioningmakespan minimization

Exploiting Application-to-Architecture Dependencies for Designing Scalable OS

Jan 02, 2025
YX
Yao Xiao
🏛️ University of Southern California | Cisco Research

Traditional operating systems suffer from poor scalability on many-core processors and low parallel efficiency due to their inability to perceive application semantics. To address this, we propose NetworkedOS—a novel application-aware, networked OS architecture. Our approach leverages compile-time dynamic instruction dependency analysis to construct a multi-layer network model that explicitly captures runtime dependencies among applications, the kernel, and hardware. We further design an overlapping graph partitioning algorithm to jointly optimize parallel execution and inter-core communication overhead, and implement a runtime process affinity mapping scheduler. Crucially, NetworkedOS breaks the conventional “black-box” OS assumption regarding application semantics for the first time. Experimental evaluation shows that NetworkedOS achieves a 7.11× speedup over Linux on a 128-core system and a 2.01× improvement over Barrelfish on a 64-core system, significantly enhancing scalability and resource utilization under large-scale parallel workloads.

Multicore ManagementParallel TasksSystem Optimization

Latest Papers

What's happening recently
View more

This work addresses the poor scalability of existing memory isolation mechanisms in supporting large-scale, fine-grained intra-process security domains by proposing an efficient linked isolation model built upon CHERI hardware capabilities. We introduce a novel "one-click" library-boundary isolation strategy, combined with deep customizations to compilers and operating systems for both Armv8-A and RISC-V architectures. This approach enables fine-grained isolation of tens of thousands of security domains and over 500 concurrent domains within UNIX user space, substantially surpassing the concurrency limitations of mechanisms such as Intel MPK. The proposed solution has been validated on Morello and Codasip X730 platforms, demonstrating that the vast majority of C/C++ programs execute without modification, requiring only minor adaptations for the V8 engine.

CHERIcompartmentalizationin-process isolation

Existing hypergraph partitioning methods often become trapped in local optima, limiting partition quality. This work proposes ComPart, a novel framework that integrates community structure guidance during both the initial partitioning and uncoarsening phases. It is the first to comprehensively incorporate community detection throughout the entire uncoarsening process and extends the theory of local dense decomposition from graphs to hypergraphs to generate high-quality initial partitions. By synergistically combining diverse community detection techniques, hypergraph local dense decomposition, and a multilevel partitioning strategy, ComPart consistently outperforms state-of-the-art methods on standard benchmarks, achieving significantly improved partition quality.

community detectionhypergraph partitioninginitial partitioning

This study addresses the coordination challenges arising from independent control across cloud, high-performance computing (HPC), and edge AI infrastructures. Conceptualizing the AI platform as a "system of systems," this work proposes an architectural paradigm characterized by usage fusion and federated control. Methodologically, it adopts a systems engineering framework that achieves cross-domain coordination through interface contracts while preserving native control planes. The approach incorporates boundary testing, responsibility models, and seven integration facets, leveraging interface mapping, policy contexts, and operational evidence to guide integration design. The primary contribution lies in establishing a unified framework for evaluating interoperability, governance capabilities, and fault isolation, thereby delineating clear directions for future research.

Cloud-HPC-Edge AI PlatformsComposable IntegrationFederated Control

This work addresses the programming challenges and error-proneness introduced by Arm’s POE2 architecture, which employs a complex spatiotemporal permission mechanism yet lacks a unified security model. We propose the first general-purpose secure programming model tailored for POE2, abstracting away the intricacies of its spatial and temporal indexing and encapsulating hardware features such as memory protection keys, dedicated registers, and table structures. By doing so, our model significantly simplifies permission management while preserving POE2’s strong security guarantees. It naturally supports common intra-process isolation patterns used in software partitioning, enabling developers to construct secure isolated systems more efficiently and with fewer errors.

architectural complexityintra-process isolationmemory protection keys

This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.

computational workflowsportablereproducible

Hot Scholars

NM

Noboru Murata

Waseda University
statistical learningmachine learning
SM

S M Hasan Mahmud

Associate Professor, Department of Software Engineering, Daffodil International University
Machine LearningImage ProcessingBioinformaticsData Science
WS

Wuzhen Shi

Shenzhen University
Image/Video Compression and EnhancementAffective ComputingAIGC
HH

Hideitsu Hino

The Institute of Statistical Mathematics
Statistical LearningMachine LearningData Mining
AS

Abhishek Sharma

Indian Institute of Technology Mandi
Statistical PhyiscsActive MatterActive GranularsActive Nematics