data collection

Designs and builds methods and systems to acquire and aggregate datasets from sensors, experiments, user interactions, logs, or external APIs, including sampling strategies, instrumentation, scraping or ingestion code, labeling protocols, metadata capture, and storage formats. Evaluates and monitors data quality, completeness, provenance, and bias during collection and organizes collected data for downstream processing and analysis.

datacollection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.65
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$206K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A methodology and a platform for high-quality rich personal data collection

Jan 28, 2025
IK
Ivan Kayongo
🏛️ University of Trento

Existing mobile sensing data collection methods neglect subjective feedback (e.g., questionnaires, self-reports), leading to fragmented contextual understanding and inaccurate behavioral modeling. To address this, we propose a human-in-the-loop collaborative sensing framework built upon the iLog platform. Our approach introduces three key innovations: (1) a dual-dimensional “context–time” modeling paradigm; (2) a calendar-style real-time monitoring dashboard; and (3) a dynamic acquisition plan revision mechanism. Leveraging context-aware modeling, real-time visual analytics, an adaptive experimental workflow engine, and purposeful human–system interaction design, the framework enhances controllability for researchers, participants, and the system itself. Evaluated with 350 university students, our method significantly improves semantic richness, contextual completeness, and overall data quality—enabling more accurate behavioral modeling and fine-grained personalized analysis.

Data CollectionSmart Device SensorsSubjective Information

This work addresses the challenge of reproducibility in actively developed experimental projects, which often suffer from unstructured data management and are overlooked by conventional data management plans. We propose a lightweight, domain-agnostic framework built upon the Sacred experiment tracking model that, from the project’s inception, systematically organizes parameters, metadata, metric trajectories, and associated files. Small-scale data are stored in a NoSQL database, while large files are linked via unique identifiers to dedicated storage systems. The framework seamlessly integrates into existing research workflows, supports both local deployment and public release, and uniquely targets the dynamic exploration phase of research. By doing so, it establishes a practical bridge from early-stage experimentation to FAIR-compliant data sharing, significantly enhancing collaborative efficiency and scientific reproducibility without compromising flexibility or scalability.

collaborative researchexperimental data managementFAIR data

The hunt for research data: Development of an open-source workflow for tracking institutionally-affiliated research data publications

Jul 01, 2025
BM
Bryan M. Gee
🏛️ University of Texas Libraries | The University of Texas at Austin

Institutions face significant challenges in systematically tracking their affiliated research data publications, primarily due to the widespread absence, inconsistency, or non-standardization of institutional attribution metadata (e.g., missing or ambiguous institutional names, lack of persistent identifiers such as DOIs) in existing data repositories. Method: We propose the first open-source, institution-centric workflow for tracking data publications, integrating over 70 open APIs and employing multi-source metadata harvesting, normalization, and automated aggregation—thereby reducing reliance on explicit attribution signals like DOIs or manually curated affiliations. Contribution/Results: The workflow enables efficient discovery and consolidation of over 4,000 cross-platform datasets. Evaluation demonstrates substantial improvements in institutional data discoverability, coverage breadth, and retrieval efficiency. This work delivers a reusable technical infrastructure to support research administration and data governance at institutional and systemic levels.

Address challenges in discovering institution-affiliated datasetsDevelop open-source workflow for tracking institutional research dataImprove metadata standardization for comprehensive data retrieval

Scientists frequently record experimental metadata in spreadsheets, yet ensuring consistency and standards compliance remains challenging. This paper introduces a spreadsheet-native metadata governance paradigm: customized Excel/CSV templates embed HuBMAP standards; OWL/SKOS ontology-driven controlled vocabularies are integrated; and a web-based real-time semantic validation tool enables immediate, on-entry verification. The approach seamlessly incorporates semantic constraints into familiar spreadsheet workflows—requiring no platform switching or new system adoption. Deployed across the HuBMAP Consortium, it significantly improved multi-omics metadata compliance rates, increased data entry efficiency, and reduced error identification and correction time by over 70%. To our knowledge, this is the first work to deeply embed ontology-based constraints and real-time semantic validation directly within spreadsheet environments, establishing a scalable, practical paradigm for biomedical metadata standardization.

Addressing spreadsheet limitations for consistent experiment-related metadata annotationEnsuring metadata standards compliance in spreadsheet-based scientific data entryProviding quality control for biomedical metadata collection using spreadsheets

Latest Papers

What's happening recently
View more

This work addresses the challenge of reconciling high throughput and low query latency in traditional ETL pipelines when processing continuously arriving fresh data, where unpredictable preprocessing operations often create bottlenecks. The authors propose Fluid ETL Pipelines, which introduce, for the first time, an elastic and non-blocking preprocessing mechanism that decouples data ingestion from transformation. By dynamically scheduling preprocessing tasks based on resource availability and user interest—without blocking data ingestion—and leveraging preemptible computing resources such as Amazon Spot instances, the approach significantly reduces operational costs. Experimental results demonstrate that Fluid ETL Pipelines substantially improve the efficiency of exploring fresh data, offering a novel direction for accelerating real-time queries and enabling adaptive preprocessing management.

data preprocessing routinesETL pipelinesfresh data exploration

Hot Scholars

SK

Sanmi Koyejo

Assistant Professor, Stanford University
Machine LearningHealthcare AINeuroinformatics
DR

Dan Roth

Professor of Computer Science, University of Pennsylvania
Natural Language ProcessingMachine LearningKnowledge Representation and ReasoningArtificial Intelligence
DS

Daniel Seita

University of Southern California
RoboticsMachine Learning
PN

Preslav Nakov

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Computational LinguisticsLarge Language ModelsFact-checkingFake News
SZ

Shanshan Zhong

Carnegie Mellon University
Language ModelsMultimodal UnderstandingMultimodal Generation