disaggregated systems design

Designs and evaluates system architectures that decompose traditionally monolithic resources into modular, networked components; specifies the hardware and software interfaces, interconnect protocols, resource orchestration, placement, consistency, performance and fault‑tolerance mechanisms needed to compose, share, and scale compute, memory, storage, and accelerator resources independently.

disaggregatedsystemsdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic guidance for enterprise software teams in choosing between monolithic and microservices architectures. The work proposes a decision-making framework that integrates technical and organizational factors, evaluating the trade-offs of each architecture across dimensions such as scalability, reliability, deployment efficiency, and organizational complexity. The assessment is grounded in system scale, business requirements, operational maturity, and long-term maintainability. Through architectural pattern analysis, a structured evaluation model, and multiple case studies, the authors develop a practical selection methodology tailored to real-world engineering contexts. This approach offers enterprises clear architectural evolution pathways and actionable guidelines aligned with their developmental stages, thereby significantly enhancing the rationality and sustainability of system design decisions.

MicroservicesMonolithic ArchitectureOrganizational Complexity

Exploiting Application-to-Architecture Dependencies for Designing Scalable OS

Jan 02, 2025
YX
Yao Xiao
🏛️ University of Southern California | Cisco Research

Traditional operating systems suffer from poor scalability on many-core processors and low parallel efficiency due to their inability to perceive application semantics. To address this, we propose NetworkedOS—a novel application-aware, networked OS architecture. Our approach leverages compile-time dynamic instruction dependency analysis to construct a multi-layer network model that explicitly captures runtime dependencies among applications, the kernel, and hardware. We further design an overlapping graph partitioning algorithm to jointly optimize parallel execution and inter-core communication overhead, and implement a runtime process affinity mapping scheduler. Crucially, NetworkedOS breaks the conventional “black-box” OS assumption regarding application semantics for the first time. Experimental evaluation shows that NetworkedOS achieves a 7.11× speedup over Linux on a 128-core system and a 2.01× improvement over Barrelfish on a 64-core system, significantly enhancing scalability and resource utilization under large-scale parallel workloads.

Multicore ManagementParallel TasksSystem Optimization

This work addresses the lack of unified and reproducible evaluation criteria for service boundary identification in the migration from monolithic systems to microservices. It presents the first systematic comparison of mainstream microservice decomposition approaches—static, dynamic, and hybrid—within a consistent experimental framework. Using standardized metric computation procedures and multiple benchmark systems (JPetStore, AcmeAir, DayTrader, Plants), the study evaluates key dimensions including structural modularity (SM), interface count (IFN), inter-component communication (ICP), and non-extreme distribution (NED). Experimental results demonstrate that HDBScan with hierarchical clustering consistently yields highly cohesive and loosely coupled microservice partitions across benchmarks, achieving an optimal trade-off between modularity strength and communication overhead, thereby significantly enhancing the objectivity and reproducibility of microservice decomposition evaluation.

benchmark evaluationmicroservice decompositionmonolithic architecture

Microservices Are Dying, A New Method for Module Division Based on Universal Interfaces

Nov 06, 2025
QW
Qing Wang
🏛️ BNRist | Tsinghua University

Microservices achieve physical isolation but fail to prevent the proliferation of logical coupling, undermining module independence. This paper proposes a novel modularization paradigm based on universal interface boundaries, constructs a quantifiable model for assessing module independence, and designs a runtime mechanism supporting dynamic loading, unloading, and hot updates within a single process. Its core contributions are: (1) reframing module independence as a formal, modelable, and measurable system property—moving beyond qualitative assertions; (2) replacing implicit dependencies with explicit interface contracts to fundamentally block coupling propagation; and (3) implementing the EIGHT platform prototype, which achieves microservice-level module autonomy within a monolithic process. Experimental results demonstrate that the approach significantly reduces the impact scope of cross-module changes, enhancing system maintainability and evolutionary efficiency. It provides both theoretical foundations and practical pathways for next-generation architectures transcending the monolith–microservice dichotomy.

Addressing dependency propagation in microservices through module independence calculationDeveloping EIGHT platform architecture enabling dynamic runtime modifications in monolithic applicationsProposing universal interfaces to eliminate inter-module dependencies in system design

This study addresses the coordination challenges arising from independent control across cloud, high-performance computing (HPC), and edge AI infrastructures. Conceptualizing the AI platform as a "system of systems," this work proposes an architectural paradigm characterized by usage fusion and federated control. Methodologically, it adopts a systems engineering framework that achieves cross-domain coordination through interface contracts while preserving native control planes. The approach incorporates boundary testing, responsibility models, and seven integration facets, leveraging interface mapping, policy contexts, and operational evidence to guide integration design. The primary contribution lies in establishing a unified framework for evaluating interoperability, governance capabilities, and fault isolation, thereby delineating clear directions for future research.

Cloud-HPC-Edge AI PlatformsComposable IntegrationFederated Control

Latest Papers

What's happening recently
View more

This work addresses the lack of systematic, reproducible, and maintainable testing methodologies in existing dynamic resource management libraries. We propose an automated validation framework tailored for high-performance computing (HPC) environments, which introduces a novel multi-level testing taxonomy encompassing both functional and non-functional requirements. Built upon an MPI-based scalable library testing methodology, the framework supports core primitives of dynamic resource management systems—such as initialization, readiness checks, and reconfiguration—and integrates containerized virtual clusters with continuous integration (CI) ecosystems. Experimental evaluation demonstrates that our approach significantly improves early fault detection rates, reduces maintenance overhead caused by evolving dependencies, and is readily generalizable to other systems exhibiting similar variability mechanisms.

dynamic resource managementHPClibrary correctness

Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.

Compound AI SystemsDistributed AIModel-Centric Design

This study addresses cascading failures caused by shared firmware in commercial multi-host network interface cards (NICs) and the operational challenges of hyperscale deployments by proposing fbnic, a system-level solution. Architecturally, it introduces physical isolation and driver-priority mechanisms, combined with sub-sled-granularity firmware upgrade orchestration and slice-level fault containment. Operationally, it establishes a hardware-in-the-loop (HIL) continuous integration pipeline alongside cross-layer fault attribution monitoring and an automated remediation toolchain. Deployed across hundreds of thousands of hosts, fbnic reduces unplanned unavailability by 12×, shortens mean time to repair by 37%, and decreases hardware replacement rates by 2.3×, demonstrating robust stability at scale.

custom hardwarehyperscalerisolation failures

This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.

computational workflowsportablereproducible

Hot Scholars

YX

Ying Xiong

Clausthal University of Technology
Petroleum geologySedimentologyGeochemistry
ZF

Zhenan Fan

Staff Researcher at Huawei Technologies Canada
OptimizationLarge Language Model
XW

Xinglu Wang

PhD student, SFU
OptimizationMulti-task learning
NG

Niloofar Gholipour

PhD Candidate @ Univ. of Quebec
Distributed systemsDeep Reinforcement Learning
GH

Guoyu Hu

National University of Singapore
searchdatabasecloud