Score
Designs and builds methods and systems to acquire and aggregate datasets from sensors, experiments, user interactions, logs, or external APIs, including sampling strategies, instrumentation, scraping or ingestion code, labeling protocols, metadata capture, and storage formats. Evaluates and monitors data quality, completeness, provenance, and bias during collection and organizes collected data for downstream processing and analysis.
Existing mobile sensing data collection methods neglect subjective feedback (e.g., questionnaires, self-reports), leading to fragmented contextual understanding and inaccurate behavioral modeling. To address this, we propose a human-in-the-loop collaborative sensing framework built upon the iLog platform. Our approach introduces three key innovations: (1) a dual-dimensional “context–time” modeling paradigm; (2) a calendar-style real-time monitoring dashboard; and (3) a dynamic acquisition plan revision mechanism. Leveraging context-aware modeling, real-time visual analytics, an adaptive experimental workflow engine, and purposeful human–system interaction design, the framework enhances controllability for researchers, participants, and the system itself. Evaluated with 350 university students, our method significantly improves semantic richness, contextual completeness, and overall data quality—enabling more accurate behavioral modeling and fine-grained personalized analysis.
This work addresses the challenge of reproducibility in actively developed experimental projects, which often suffer from unstructured data management and are overlooked by conventional data management plans. We propose a lightweight, domain-agnostic framework built upon the Sacred experiment tracking model that, from the project’s inception, systematically organizes parameters, metadata, metric trajectories, and associated files. Small-scale data are stored in a NoSQL database, while large files are linked via unique identifiers to dedicated storage systems. The framework seamlessly integrates into existing research workflows, supports both local deployment and public release, and uniquely targets the dynamic exploration phase of research. By doing so, it establishes a practical bridge from early-stage experimentation to FAIR-compliant data sharing, significantly enhancing collaborative efficiency and scientific reproducibility without compromising flexibility or scalability.
Institutions face significant challenges in systematically tracking their affiliated research data publications, primarily due to the widespread absence, inconsistency, or non-standardization of institutional attribution metadata (e.g., missing or ambiguous institutional names, lack of persistent identifiers such as DOIs) in existing data repositories. Method: We propose the first open-source, institution-centric workflow for tracking data publications, integrating over 70 open APIs and employing multi-source metadata harvesting, normalization, and automated aggregation—thereby reducing reliance on explicit attribution signals like DOIs or manually curated affiliations. Contribution/Results: The workflow enables efficient discovery and consolidation of over 4,000 cross-platform datasets. Evaluation demonstrates substantial improvements in institutional data discoverability, coverage breadth, and retrieval efficiency. This work delivers a reusable technical infrastructure to support research administration and data governance at institutional and systemic levels.
Scientists frequently record experimental metadata in spreadsheets, yet ensuring consistency and standards compliance remains challenging. This paper introduces a spreadsheet-native metadata governance paradigm: customized Excel/CSV templates embed HuBMAP standards; OWL/SKOS ontology-driven controlled vocabularies are integrated; and a web-based real-time semantic validation tool enables immediate, on-entry verification. The approach seamlessly incorporates semantic constraints into familiar spreadsheet workflows—requiring no platform switching or new system adoption. Deployed across the HuBMAP Consortium, it significantly improved multi-omics metadata compliance rates, increased data entry efficiency, and reduced error identification and correction time by over 70%. To our knowledge, this is the first work to deeply embed ontology-based constraints and real-time semantic validation directly within spreadsheet environments, establishing a scalable, practical paradigm for biomedical metadata standardization.
This work addresses the challenge of reconciling high throughput and low query latency in traditional ETL pipelines when processing continuously arriving fresh data, where unpredictable preprocessing operations often create bottlenecks. The authors propose Fluid ETL Pipelines, which introduce, for the first time, an elastic and non-blocking preprocessing mechanism that decouples data ingestion from transformation. By dynamically scheduling preprocessing tasks based on resource availability and user interest—without blocking data ingestion—and leveraging preemptible computing resources such as Amazon Spot instances, the approach significantly reduces operational costs. Experimental results demonstrate that Fluid ETL Pipelines substantially improve the efficiency of exploring fresh data, offering a novel direction for accelerating real-time queries and enabling adaptive preprocessing management.