data acquisition

Designs and implements systems and processes to collect and ingest data from diverse sources into storage or processing pipelines. This includes building instrumentation and logging, connectors and APIs, scraping or export tools, and ingestion/ETL workflows with validation and metadata capture to ensure data fidelity and availability.

dataacquisition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.94
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$225K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Nagare Media Ingest: A System for Multimedia Ingest Workflows

Sep 15, 2025
MN
Matthias Neugebauer
🏛️ University of Münster

To address the challenges of protocol heterogeneity, architectural rigidity, and limited scalability in cloud-edge collaborative multimedia data ingestion, this paper proposes and implements an open-source, modular multimedia ingestion system. The system innovatively decouples the ingestion pipeline into configurable, concurrently executing microservices, supporting mainstream streaming protocols—including SRT, RIST, DASH-IF LMI, and MOQT—and is designed following cloud-native principles to ensure compatibility with Kubernetes as well as lightweight edge deployments. Compared to monolithic ingestion solutions, it significantly enhances protocol adaptability, workflow customizability, and horizontal scalability. Real-world deployment evaluations demonstrate its robust stability and high resource efficiency under high-concurrency, multi-source heterogeneous scenarios. The system provides a reusable, production-ready infrastructure foundation for modern multimedia workflows.

Handling complexity of multimedia ingest workflowsProviding flexible component-based ingest system designSupporting diverse streaming protocols and use cases

This work addresses the challenge of reconciling high throughput and low query latency in traditional ETL pipelines when processing continuously arriving fresh data, where unpredictable preprocessing operations often create bottlenecks. The authors propose Fluid ETL Pipelines, which introduce, for the first time, an elastic and non-blocking preprocessing mechanism that decouples data ingestion from transformation. By dynamically scheduling preprocessing tasks based on resource availability and user interest—without blocking data ingestion—and leveraging preemptible computing resources such as Amazon Spot instances, the approach significantly reduces operational costs. Experimental results demonstrate that Fluid ETL Pipelines substantially improve the efficiency of exploring fresh data, offering a novel direction for accelerating real-time queries and enabling adaptive preprocessing management.

data preprocessing routinesETL pipelinesfresh data exploration

FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Jul 30, 2025
MD
Mattia Di Profio
🏛️ University of Aberdeen

Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.

Automating ETL workflows to reduce human interventionDesigning context-specific transformations without manual inputStandardizing diverse datasets using example-driven approaches

PRE-Share Data: Assistance Tool for Resource-aware Designing of Data-sharing Pipelines

Mar 17, 2025
SM
Sepideh Masoudi
🏛️ Technische Universität Berlin

In cross-organizational data sharing, existing multi-pipeline transformation design suffers from low efficiency and severe resource waste under dual constraints of governance compliance and recipient-side adaptability. Method: This paper proposes a reuse-aware pipeline design assistance paradigm that integrates flowchart-based modeling, semantic matching of transformation operations, fine-grained resource consumption modeling, and heuristic configuration optimization. It enables automatic identification of reusable transformation components across pipelines, recommends optimal pipeline structures, and quantifies potential resource savings. Contribution/Results: As the first design assistance framework supporting predictive reporting generation, it achieves, on real-world use cases, an average 37% reduction in computational resource consumption and a 52% reduction in design cycle time, while remaining compatible with self-service data platform deployments.

Designing efficient data-sharing pipelines across organizationsEnsuring compliance with governance policies and recipient requirementsReusing transformation processes to optimize resource consumption

Applying Process Mining on Scientific Workflows: a Case Study

Jul 06, 2023
ZS
Zahra Sadeghibogar
🏛️ RWTH Aachen University

SLURM logs in HPC scientific workflows lack explicit case identifiers, hindering direct application of process mining. Method: This paper proposes an automatic job-correlation method based on implicit job dependency modeling—parsing SLURM logs and jointly leveraging spatiotemporal job feature matching and graph-structured modeling to achieve end-to-end clustering of unannotated jobs. Contribution/Results: We introduce the first systematic preprocessing framework for process mining on HPC logs, integrating algorithms such as Heuristics Miner to support process discovery and bottleneck diagnosis. Evaluated on real-world HPC cluster logs, our approach significantly improves workflow traceability, accurately identifies I/O- and scheduler-related performance bottlenecks, and enables high-fidelity reconstruction of end-to-end process models.

Correlate jobs with explicit or implicit dependencies.Document workflows and identify performance bottlenecks.Extract case IDs from SLURM-based HPC logs.

Latest Papers

What's happening recently
View more

This work addresses the challenge of reproducibility in actively developed experimental projects, which often suffer from unstructured data management and are overlooked by conventional data management plans. We propose a lightweight, domain-agnostic framework built upon the Sacred experiment tracking model that, from the project’s inception, systematically organizes parameters, metadata, metric trajectories, and associated files. Small-scale data are stored in a NoSQL database, while large files are linked via unique identifiers to dedicated storage systems. The framework seamlessly integrates into existing research workflows, supports both local deployment and public release, and uniquely targets the dynamic exploration phase of research. By doing so, it establishes a practical bridge from early-stage experimentation to FAIR-compliant data sharing, significantly enhancing collaborative efficiency and scientific reproducibility without compromising flexibility or scalability.

collaborative researchexperimental data managementFAIR data

This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.

Data ReliabilityDeveloper ProductivityDORA Metrics

In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.

Enhances auditability and traceability in regulated data pipelinesIntegrates data quality control into continuous DataOps managementUnifies rule-based, statistical, and AI methods for anomaly detection

Hot Scholars

SV

Sergei V. Kalinin

Weston Fulton Chair Professor, UT Knoxville. Chief Scientist, AI/ML for Physical Sciences, PNNL
AI4Materialsautomated experimentelectron microscopySPM
YS

Yitian Shao

HITSZ, TU Dresden, MPI-IS, UCSB
HapticsRoboticsWearable ElectronicsAR/VR
HX

He Xu

Nanjing University of Posts and Telecommunications
IoT
YQ

Yi Qin

Chongqing University
signal processingfault diagnosisartificial intelligencemeasurement
JX

Junshi Xia

RIKEN AIP
Machine LearningClassificationRemote Sensing