Score
Designs, builds, and operates scalable systems that collect, ingest, store, process, and serve data, including pipelines, storage and compute layers, metadata/catalogs, APIs, and monitoring/deployment components. Implements data governance, security, reliability, performance tuning, and integration with analytics and downstream consumers.
In the context of open science, there is an urgent need to implement the FAIR principles (Findable, Accessible, Interoperable, Reusable) in scientific data management, yet systematic guidance on achieving FAIR compliance through big data software reference architectures (SRAs) remains lacking. Method: We conducted a rigorous systematic literature review, screening 323 publications—including those from authoritative databases and expert recommendations—and performed structured data extraction and evaluation aligned with predefined research questions. Contribution/Results: The study identifies seven generic FAIR-compliant SRAs, thirteen scenario-specific FAIR pipelines, and three fully FAIR-compatible SRAs. It uncovers critical bottlenecks in metadata standardization, cross-platform interoperability, and long-term reusability. Furthermore, it establishes the first classification framework and empirical evaluation system for FAIR-oriented big data SRAs, thereby filling a significant research gap and providing a methodological foundation and strategic direction for future SRA design, policy formulation, and tool development.
This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.
This study addresses the lack of systematic optimization in cloud data pipelines with respect to cost, execution time, and resource utilization, particularly in multi-tenant and industrial settings where research remains limited. Through a comprehensive systematic literature review, the work establishes a unified classification framework for optimization objectives that encompasses both single- and multi-cloud environments as well as batch and stream processing paradigms. The analysis synthesizes existing approaches and identifies critical research gaps, including insufficient support for multi-tenancy, inadequate multi-cloud coordination, and a scarcity of real-world deployment validation. By clarifying the core objectives and technical pathways for optimizing cloud data pipelines, this paper provides a theoretical foundation and clear direction for future research in this domain.
Existing data systems lack embedded accountability mechanisms when supporting decision-making, yielding outputs that, while efficient and accurate, often fail to guarantee justifiability, constraint compliance, and operational feasibility. This work proposes RAIDS, a novel framework that treats accountability not as post-hoc metadata but as an integral execution semantics. RAIDS introduces responsibility contracts as a first-class operator abstraction and enforces end-to-end accountability through responsibility state propagation and preservation mechanisms spanning the entire pipeline from data processing to decision generation. The framework formally defines responsibility-preserving objectives and establishes a full-stack infrastructure for accountable intelligence, encompassing execution, optimization, provenance, and evaluation. Furthermore, it articulates a comprehensive research agenda for responsibility-aware data systems, laying the theoretical foundation for decision systems that are trustworthy, controllable, and auditable.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
This study addresses the growing complexity of operating large-scale computing infrastructures, such as high-performance computing (HPC) systems, where existing operational data analytics (ODA) frameworks struggle to effectively support multi-layered, distributed graph processing ecosystems. The work provides a systematic review of core ODA components, benchmarks prevailing approaches, and proposes a novel holistic ODA framework that integrates the distributed graph processing hierarchy introduced by Sherif Sak et al., thereby extending the functional scope of Netti et al.’s prior work. By unifying fine-grained monitoring, ODA architecture, and graph processing system design, the proposed framework significantly enhances structural integrity and functional extensibility, markedly improving operational efficiency. Furthermore, it illuminates key research directions for ODA in high-performance computing environments.
To address insufficient multi-objective load balancing in stream processing systems under complex workloads, this paper proposes a multi-tier collaborative scheduling framework. The framework introduces dynamic inter-layer coordination mechanisms and lightweight interfaces among schedulers, enabling seamless integration of novel scheduling policies. It jointly optimizes computational resource utilization, end-to-end latency, and throughput by integrating multi-objective optimization, distributed resource management, and real-time feedback control. Its key innovation lies in shifting hierarchical scheduling from static decoupling to dynamic collaboration—preserving scalability while significantly enhancing adaptability. Evaluated in Meta’s production environment, the system reliably processes TB-scale data with sub-second latency; it improves critical resource utilization by 27% and reduces tail latency by 41%.
This work proposes a hierarchical meta-agent architecture to overcome the static and rigid nature of traditional data processing pipelines, which struggle to autonomously monitor and optimize themselves post-deployment. The framework integrates three core components—planning, agent orchestration, and a monitoring feedback loop—to enable dynamic construction, execution, and iterative refinement of end-to-end workflows. It introduces context-aware optimization, adaptive workload partitioning, and progressive sampling mechanisms, facilitating agent reuse and seamless integration with external tools. These innovations significantly enhance system flexibility and scalability. Experimental results demonstrate the framework’s effectiveness and practicality in automatically constructing, continuously monitoring, and adaptively optimizing diverse data processing tasks.
In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.
This work addresses the limitation of existing distributed data pipeline systems, which require users to explicitly define complete workflow graphs, by proposing a unified planning and scheduling framework that automatically constructs end-to-end persistent pipelines from implicit goal declarations alone. The approach introduces, for the first time, a numeric-domain-independent planner into the context of persistent scheduling, integrating workflow and resource graph modeling, numeric planning, and network interface scheduling to achieve full automation. Experimental results demonstrate the feasibility and scalability of the method: under a single-machine constraint of one hour of CPU time and 30 GB of memory, the system successfully scheduled a linear pipeline spanning eight sites and comprising fourteen components.