Score
Designs, builds, and analyzes systems, pipelines, and models that store, process, and extract insight from datasets whose size, dimensionality, or throughput require engineering beyond single‑machine in‑memory tools. Work includes scalable data ingestion and storage, partitioning and indexing, distributed or parallel processing, sampling and summarization strategies, and algorithms for cleaning, querying, and modeling large‑scale data.
To address the lack of unified benchmarks for model performance evaluation on high-dimensional big data in both local and distributed environments, this work designs an end-to-end evaluation framework covering three representative tasks—Epsilon (numerical regression), RestMex (text classification), and IMDb (movie feature analysis). Leveraging Apache Spark (Scala), we establish a reproducible heterogeneous computing experimental infrastructure to systematically compare traditional machine learning and deep learning models across accuracy, training efficiency, and resource consumption. This study presents the first pedagogically implemented standardized benchmark supporting multiple models, multimodal data, and diverse deployment scenarios, empirically uncovering performance bottlenecks and architectural trade-offs inherent in distributed scaling. The outcomes include an open-source evaluation pipeline, a standardized reporting template, and a reusable teaching paradigm—providing empirical foundations for AI system selection and optimization in big data contexts.
AI-driven ML workloads on HPC systems exhibit a novel I/O pattern—characterized by massive small-file random reads—that diverges significantly from traditional HPC applications, causing severe performance bottlenecks in parallel file systems (e.g., Lustre, GPFS). Method: Based on empirical studies conducted from 2019–2024 and bibliometric analysis of 300+ publications, we develop the first comprehensive ML-HPC I/O analytical framework. Leveraging I/O profiling tools (IOtracer, Darshan, LMT) and runtime logs from PyTorch/TensorFlow, we systematically characterize I/O behavior across data preprocessing, training, and inference stages. Contribution/Results: We identify six critical research gaps and propose I/O-aware ML-system co-design principles. This work establishes a foundational theoretical framework and practical guidelines for designing AI-ready HPC storage architectures, bridging the gap between ML workload requirements and HPC I/O system capabilities.
To address the joint optimization challenges of performance, maintainability, and collaborative efficiency in large-scale integrated machine learning within distributed data processing systems, this paper proposes Pipes—a declarative, modular data pipeline architecture. Pipes decomposes pipelines into logically encapsulated computation units, implemented atop Apache Spark with standardized interfaces and well-defined component boundaries—departing from conventional microservice paradigms to enable high-performance, maintainable ML pipeline development. In enterprise deployments, Pipes improves development efficiency by 50%, reduces collaborative debugging cycles from weeks to days, achieves 500× scalability, and delivers 10× higher throughput. Academic benchmarks show >5.7× throughput improvement and 99% CPU utilization. Its core contribution is the first deep integration of declarative abstractions with Spark’s native execution model, simultaneously advancing both development methodology and system performance.
This work addresses the challenge of reconciling high throughput and low query latency in traditional ETL pipelines when processing continuously arriving fresh data, where unpredictable preprocessing operations often create bottlenecks. The authors propose Fluid ETL Pipelines, which introduce, for the first time, an elastic and non-blocking preprocessing mechanism that decouples data ingestion from transformation. By dynamically scheduling preprocessing tasks based on resource availability and user interest—without blocking data ingestion—and leveraging preemptible computing resources such as Amazon Spot instances, the approach significantly reduces operational costs. Experimental results demonstrate that Fluid ETL Pipelines substantially improve the efficiency of exploring fresh data, offering a novel direction for accelerating real-time queries and enabling adaptive preprocessing management.
This work addresses memory exhaustion and I/O bottlenecks in processing petabyte-scale image datasets—such as 1.4 PB electron microscopy volumes or 150 TB organ atlases—by introducing a streaming single-pass architecture based on a sweep execution model. The approach aligns disk reads with a one-dimensional sweep order and combines windowed operations with overlap-aware tiling to enable efficient processing under tight memory constraints. A domain-specific language (DSL) is designed to automatically optimize window sizes, fuse pipeline stages, and schedule multi-pass sweeps at compile time and runtime. The system supports Zarr, HDF5, and slice-based formats without requiring full-image residency in memory, achieving significantly higher throughput, near-linear I/O scaling, and predictable memory usage while seamlessly integrating with existing segmentation and morphological analysis toolchains.
Accurately estimating resource requirements for scientific workflows remains challenging due to diverse analytical scenarios, varying user expertise, and highly heterogeneous computing platforms. To address this, we propose an end-to-end machine learning framework that directly learns CPU, memory, and runtime requirements for each workflow step from historical task execution profiles—eliminating reliance on domain-specific heuristics or time-consuming two-stage trial runs. Integrated into the PanDA workflow management system, our framework enables proactive, fine-grained dynamic resource pre-allocation. Experimental evaluation in large-scale scientific computing environments—including the Large Hadron Collider (LHC)—demonstrates that our approach significantly outperforms baseline methods: average resource waste is reduced by 32%, scheduling latency decreases by 27%, and heterogeneous resource utilization and workflow execution stability are substantially improved.
To address systemic inefficiencies—including resource idleness, redundant data transfers, and violations of data locality—arising from misaligned coordination between the PanDA workflow system and the Rucio data management system in the ATLAS experiment, this paper proposes an end-to-end co-optimization framework. We introduce a novel file-level metadata matching algorithm to precisely associate computing tasks with datasets, and integrate log-based tracing, spatiotemporal imbalance analysis, and anomaly pattern detection to construct a fine-grained, holistic view of data access and movement. Our approach is the first to identify, in production, the root causes of cross-system scheduling mismatches, delivering interpretable performance insights. Empirical validation confirms tangible improvements in resource utilization and system resilience, demonstrating the feasibility and effectiveness of the proposed co-design strategies.
To address the high communication overhead of multilingual persistent operations and the suboptimal cross-model processing efficiency caused by monolithic storage engines in existing multimodel databases, this paper proposes an integrated multimodel storage engine architecture. The architecture unifies heterogeneous storage engines, each specialized and optimized for a distinct data model; introduces a multi-stage hash join algorithm to enable efficient cross-model joins; and implements unified query plan compilation and coordinated execution across models. Experimental evaluation demonstrates that the system achieves up to 188× speedup over the best-performing baseline on representative multimodel analytical workloads, while significantly improving both performance and scalability under complex, mixed-model query loads.
Scheduling data-intensive workloads in large-scale distributed systems faces challenges including complexity, heterogeneous parallelism, data locality constraints, and multi-dimensional QoS optimization (e.g., timeliness, fault tolerance, energy efficiency). Method: This paper proposes a novel workload classification scheme grounded in data characteristics and service requirements; systematically surveys and structures mainstream scheduling strategies, exposing critical limitations in dynamic adaptability, fine-grained fault tolerance, and energy–QoS co-optimization; and introduces a unified scheduling framework integrating data-locality awareness, elastic parallel scheduling, QoS-tiered guarantees, and energy-aware resource allocation. Contribution/Results: The study establishes a scalable classification paradigm, delivers a clear technology evolution roadmap, and identifies a prioritized list of open research challenges—thereby advancing foundational understanding and guiding future design of intelligent, holistic schedulers for modern distributed data systems.
This work addresses the high latency introduced by traditional decompression in scientific data analysis, which undermines the storage and transmission benefits of compression. To overcome this limitation, the authors propose a multi-stage, error-bounded decompression and homomorphic analysis framework. By abstracting a generic compression pipeline, the framework enables hierarchical partial decompression and introduces homomorphic operation algorithms tailored to three representative scientific analysis tasks, allowing computations to be performed directly on intermediate compressed representations without full decompression. Implemented atop four mainstream compressors and evaluated across five real-world datasets, the approach consistently reduces data access latency and significantly improves analytical efficiency across diverse workloads.