large datasets

Designs, builds, and analyzes systems, pipelines, and models that store, process, and extract insight from datasets whose size, dimensionality, or throughput require engineering beyond single‑machine in‑memory tools. Work includes scalable data ingestion and storage, partitioning and indexing, distributed or parallel processing, sampling and summarization strategies, and algorithms for cleaning, querying, and modeling large‑scale data.

largedatasets

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$188K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments

Dec 11, 2025
JJ
José Julián Rodríguez Gutiérrez
🏛️ División de Ingenierías Campus Irapuato-Salamanca

To address the lack of unified benchmarks for model performance evaluation on high-dimensional big data in both local and distributed environments, this work designs an end-to-end evaluation framework covering three representative tasks—Epsilon (numerical regression), RestMex (text classification), and IMDb (movie feature analysis). Leveraging Apache Spark (Scala), we establish a reproducible heterogeneous computing experimental infrastructure to systematically compare traditional machine learning and deep learning models across accuracy, training efficiency, and resource consumption. This study presents the first pedagogically implemented standardized benchmark supporting multiple models, multimodal data, and diverse deployment scenarios, empirically uncovering performance bottlenecks and architectural trade-offs inherent in distributed scaling. The outcomes include an open-source evaluation pipeline, a standardized reporting template, and a reusable teaching paradigm—providing empirical foundations for AI system selection and optimization in big data contexts.

Benchmark machine learning architectures for high-dimensional data processingCompare local and distributed computing environments for big dataImplement workflows for text analysis and classification tasks

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Apr 16, 2024
NL
Noah Lewis
🏛️ Louisiana State University | Lawrence Berkeley National Laboratory | The Ohio State University

AI-driven ML workloads on HPC systems exhibit a novel I/O pattern—characterized by massive small-file random reads—that diverges significantly from traditional HPC applications, causing severe performance bottlenecks in parallel file systems (e.g., Lustre, GPFS). Method: Based on empirical studies conducted from 2019–2024 and bibliometric analysis of 300+ publications, we develop the first comprehensive ML-HPC I/O analytical framework. Leveraging I/O profiling tools (IOtracer, Darshan, LMT) and runtime logs from PyTorch/TensorFlow, we systematically characterize I/O behavior across data preprocessing, training, and inference stages. Contribution/Results: We identify six critical research gaps and propose I/O-aware ML-system co-design principles. This work establishes a foundational theoretical framework and practical guidelines for designing AI-ready HPC storage architectures, bridging the gap between ML workload requirements and HPC I/O system capabilities.

Analyzes I/O challenges in ML on HPC systemsExplores I/O patterns in ML training and inferenceIdentifies research gaps for future ML I/O optimization

Declarative Data Pipeline for Large Scale ML Services

Aug 20, 2025
YY
Yunzhao Yang
🏛️ Amazon Web Services

To address the joint optimization challenges of performance, maintainability, and collaborative efficiency in large-scale integrated machine learning within distributed data processing systems, this paper proposes Pipes—a declarative, modular data pipeline architecture. Pipes decomposes pipelines into logically encapsulated computation units, implemented atop Apache Spark with standardized interfaces and well-defined component boundaries—departing from conventional microservice paradigms to enable high-performance, maintainable ML pipeline development. In enterprise deployments, Pipes improves development efficiency by 50%, reduces collaborative debugging cycles from weeks to days, achieves 500× scalability, and delivers 10× higher throughput. Academic benchmarks show >5.7× throughput improvement and 99% CPU utilization. Its core contribution is the first deep integration of declarative abstractions with Spark’s native execution model, simultaneously advancing both development methodology and system performance.

Balancing performance with maintainability in large-scale ML servicesIntegrating machine learning capabilities efficiently within Apache SparkReducing communication overhead in collaborative data processing environments

This work addresses the challenge of reconciling high throughput and low query latency in traditional ETL pipelines when processing continuously arriving fresh data, where unpredictable preprocessing operations often create bottlenecks. The authors propose Fluid ETL Pipelines, which introduce, for the first time, an elastic and non-blocking preprocessing mechanism that decouples data ingestion from transformation. By dynamically scheduling preprocessing tasks based on resource availability and user interest—without blocking data ingestion—and leveraging preemptible computing resources such as Amazon Spot instances, the approach significantly reduces operational costs. Experimental results demonstrate that Fluid ETL Pipelines substantially improve the efficiency of exploring fresh data, offering a novel direction for accelerating real-time queries and enabling adaptive preprocessing management.

data preprocessing routinesETL pipelinesfresh data exploration

This work addresses memory exhaustion and I/O bottlenecks in processing petabyte-scale image datasets—such as 1.4 PB electron microscopy volumes or 150 TB organ atlases—by introducing a streaming single-pass architecture based on a sweep execution model. The approach aligns disk reads with a one-dimensional sweep order and combines windowed operations with overlap-aware tiling to enable efficient processing under tight memory constraints. A domain-specific language (DSL) is designed to automatically optimize window sizes, fuse pipeline stages, and schedule multi-pass sweeps at compile time and runtime. The system supports Zarr, HDF5, and slice-based formats without requiring full-image residency in memory, achieving significantly higher throughput, near-linear I/O scaling, and predictable memory usage while seamlessly integrating with existing segmentation and morphological analysis toolchains.

I/O-boundimage processinglarger-than-memory

Latest Papers

What's happening recently
View more

Machine Learning-Driven Predictive Resource Management in Complex Science Workflows

Sep 14, 2025
TC
Tasnuva Chowdhury
🏛️ Brookhaven National Laboratory | University of Massachusetts, Amherst | University of Pittsburgh | Carnegie Mellon University | Oak Ridge National Laboratory | SLAC National Accelerator Laboratory

Accurately estimating resource requirements for scientific workflows remains challenging due to diverse analytical scenarios, varying user expertise, and highly heterogeneous computing platforms. To address this, we propose an end-to-end machine learning framework that directly learns CPU, memory, and runtime requirements for each workflow step from historical task execution profiles—eliminating reliance on domain-specific heuristics or time-consuming two-stage trial runs. Integrated into the PanDA workflow management system, our framework enables proactive, fine-grained dynamic resource pre-allocation. Experimental evaluation in large-scale scientific computing environments—including the Large Hadron Collider (LHC)—demonstrates that our approach significantly outperforms baseline methods: average resource waste is reduced by 32%, scheduling latency decreases by 27%, and heterogeneous resource utilization and workflow execution stability are substantially improved.

Enabling optimal resource allocation using machine learningOvercoming inaccurate initial resource estimation challengesPredicting resource needs for complex scientific workflows

Data Management System Analysis for Distributed Computing Workloads

Oct 01, 2025
KH
Kuan-Chieh Hsu
🏛️ Brookhaven National Laboratory | University of Pittsburgh | Carnegie Mellon University | Oak Ridge National Laboratory | University of Massachusetts | SLAC National Accelerator Laboratory

To address systemic inefficiencies—including resource idleness, redundant data transfers, and violations of data locality—arising from misaligned coordination between the PanDA workflow system and the Rucio data management system in the ATLAS experiment, this paper proposes an end-to-end co-optimization framework. We introduce a novel file-level metadata matching algorithm to precisely associate computing tasks with datasets, and integrate log-based tracing, spatiotemporal imbalance analysis, and anomaly pattern detection to construct a fine-grained, holistic view of data access and movement. Our approach is the first to identify, in production, the root causes of cross-system scheduling mismatches, delivering interpretable performance insights. Empirical validation confirms tangible improvements in resource utilization and system resilience, demonstrating the feasibility and effectiveness of the proposed co-design strategies.

Diagnosing systemic inefficiencies in globally distributed data workflowsLinking workflow and data management systems for performance awarenessReducing unnecessary data transfers through coordinated system optimization

M2: An Analytic System with Specialized Storage Engines for Multi-Model Workloads

Aug 04, 2025
KK
Kyoseung Koo
🏛️ Seoul National University

To address the high communication overhead of multilingual persistent operations and the suboptimal cross-model processing efficiency caused by monolithic storage engines in existing multimodel databases, this paper proposes an integrated multimodel storage engine architecture. The architecture unifies heterogeneous storage engines, each specialized and optimized for a distinct data model; introduces a multi-stage hash join algorithm to enable efficient cross-model joins; and implements unified query plan compilation and coordinated execution across models. Experimental evaluation demonstrates that the system achieves up to 188× speedup over the best-performing baseline on representative multimodel analytical workloads, while significantly improving both performance and scalability under complex, mixed-model query loads.

Handling multiple data models efficiently in analyticsOptimizing storage engines for diverse data modelsReducing communication costs in polyglot persistence systems

Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges

Oct 29, 2025
GL
Georgios L. Stavrinides
🏛️ Aristotle University of Thessaloniki

Scheduling data-intensive workloads in large-scale distributed systems faces challenges including complexity, heterogeneous parallelism, data locality constraints, and multi-dimensional QoS optimization (e.g., timeliness, fault tolerance, energy efficiency). Method: This paper proposes a novel workload classification scheme grounded in data characteristics and service requirements; systematically surveys and structures mainstream scheduling strategies, exposing critical limitations in dynamic adaptability, fine-grained fault tolerance, and energy–QoS co-optimization; and introduces a unified scheduling framework integrating data-locality awareness, elastic parallel scheduling, QoS-tiered guarantees, and energy-aware resource allocation. Contribution/Results: The study establishes a scalable classification paradigm, delivers a clear technology evolution roadmap, and identifies a prioritized list of open research challenges—thereby advancing foundational understanding and guiding future design of intelligent, holistic schedulers for modern distributed data systems.

Addressing data locality and parallelism for data-intensive applicationsMeeting QoS requirements like time constraints and energy efficiencyScheduling complex workloads in large-scale distributed systems

This work addresses the high latency introduced by traditional decompression in scientific data analysis, which undermines the storage and transmission benefits of compression. To overcome this limitation, the authors propose a multi-stage, error-bounded decompression and homomorphic analysis framework. By abstracting a generic compression pipeline, the framework enables hierarchical partial decompression and introduces homomorphic operation algorithms tailored to three representative scientific analysis tasks, allowing computations to be performed directly on intermediate compressed representations without full decompression. Implemented atop four mainstream compressors and evaluated across five real-world datasets, the approach consistently reduces data access latency and significantly improves analytical efficiency across diverse workloads.

access latencyanalytical operationsdata decompression