data collection and validation

Designs and implements systems and workflows to collect, process, and validate transcripts, including methods for transcript capture, cleaning, normalization, formatting, and metadata extraction; builds tooling and procedures to verify transcript accuracy, completeness, and conformity to schema. Analyzes transcripts to detect and annotate errors or inconsistencies and to produce validated, publication- or downstream-ready transcript datasets.

datacollectionandvalidation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Halving transcription time: A fast, user-friendly and GDPR-compliant workflow to create AI-assisted transcripts for content analysis

Mar 17, 2025
JS
Jakob Sponholz
🏛️ University of Cologne | University of Koblenz | University of Münster

Qualitative research faces significant bottlenecks in manual transcription—low efficiency, high time cost, and challenges in GDPR compliance and software interoperability. This study proposes an end-to-end AI-assisted transcription workflow: it employs a localized automatic speech recognition (ASR) model, augmented with customized text post-processing and structured format conversion modules. Crucially, it achieves native, seamless integration with leading qualitative data analysis software—including ATLAS.ti and MAXQDA—for direct import of transcripts. The system incorporates built-in GDPR compliance via fully offline processing and includes phonetic and linguistic adaptations for non-native speaker speech. Empirical evaluation across 12 real-world interviews demonstrates a 46.2% average reduction in transcription time, while preserving transcription accuracy, data privacy, and cross-platform compatibility. This workflow significantly enhances the efficiency and rigor of qualitative data preparation.

Ensures AI-generated transcripts are compatible with content analysis software.Provides GDPR-compliant, offline transcription for sensitive data handling.Reduces labor-intensive transcription time in qualitative research.

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

Generative Goal Modeling

Aug 31, 2025
AS
Ateeq Sharfuddin
🏛️ Carnegie Mellon University

In software engineering, manual extraction of goals from stakeholder interviews and subsequent goal modeling suffer from low efficiency and poor reproducibility. To address this, we propose the first end-to-end automated goal modeling method integrating textual entailment reasoning with a large language model (GPT-4o). Our approach directly generates structured goal models from unstructured interview transcripts, supporting high-level goal-to-software-operation mapping, requirement refinement, and conflict/obstacle analysis, while enabling goal provenance tracing and refinement relation inference. Evaluated on 15 cross-domain interview datasets, it achieves a goal matching rate of 62.0% (comparable to human performance), a provenance tracing accuracy of 98.7%, and a refinement relation generation accuracy of 72.2%. The core innovation lies in the first application of textual entailment to goal modeling, significantly enhancing both the accuracy and interpretability of automated goal modeling.

Constructing goal models using textual entailment techniquesEvaluating GPT-4o's accuracy in goal identification and tracingExtracting goals from interview transcripts automatically

Current agent benchmarks rely on manual auditing, which struggles to scale and often fails to identify validity flaws, thereby undermining the credibility of model capability evaluations. This work proposes the first automated AI scanner tailored for agent benchmarking, leveraging large language models and structured scoring rules to detect four categories of validity issues in agent transcripts. The system is calibrated through human annotation validation and cross-benchmark evaluation. Experiments across five prominent benchmarks uncover multiple quality issues that evade manual spot-checking, demonstrating that the proposed method effectively enables systematic auditing of benchmarks. This approach establishes a new paradigm for enhancing the reliability of agent evaluations.

agentic benchmarksautomated transcript analysisbenchmark auditing

Scientists frequently record experimental metadata in spreadsheets, yet ensuring consistency and standards compliance remains challenging. This paper introduces a spreadsheet-native metadata governance paradigm: customized Excel/CSV templates embed HuBMAP standards; OWL/SKOS ontology-driven controlled vocabularies are integrated; and a web-based real-time semantic validation tool enables immediate, on-entry verification. The approach seamlessly incorporates semantic constraints into familiar spreadsheet workflows—requiring no platform switching or new system adoption. Deployed across the HuBMAP Consortium, it significantly improved multi-omics metadata compliance rates, increased data entry efficiency, and reduced error identification and correction time by over 70%. To our knowledge, this is the first work to deeply embed ontology-based constraints and real-time semantic validation directly within spreadsheet environments, establishing a scalable, practical paradigm for biomedical metadata standardization.

Addressing spreadsheet limitations for consistent experiment-related metadata annotationEnsuring metadata standards compliance in spreadsheet-based scientific data entryProviding quality control for biomedical metadata collection using spreadsheets

Latest Papers

What's happening recently
View more

This work addresses the frequent failure of natural language processing (NLP) projects in clinical settings, which often stems from a lack of systematic engineering practices and an overemphasis on algorithms at the expense of development rigor. To bridge this gap, the paper proposes a structured methodology grounded in the Systems Development Life Cycle (SDLC) framework to guide the end-to-end construction of NLP systems for extracting clinical information from electronic health records. By integrating SDLC principles throughout the NLP development pipeline, the approach counteracts the algorithm-centric bias prevalent in conventional tutorials and establishes a reproducible, generalizable development paradigm. This systematic integration enhances both the success rate and reliability of clinical text information extraction initiatives, offering a robust foundation for real-world deployment.

Clinical Data ExtractionNatural Language ProcessingNLP System Development

Existing approaches to automatic document formatting suffer from imprecise target localization and redundant content re-reading in content-aware scenarios, compounded by the absence of a dedicated evaluation benchmark. To address these limitations, this work introduces DocFormBench—the first comprehensive evaluation benchmark specifically designed for content-aware document formatting—and proposes DocFormFlow, a decoupled workflow that separates the task into two distinct phases: “what to format” (target localization) and “how to format” (format execution). By integrating large language models with multimodal models, DocFormFlow demonstrates significant improvements in formatting accuracy and substantially reduces token consumption across multiple mainstream models, underscoring precise target localization as a critical factor for high performance.

content-awaredocument formattingevaluation benchmark

This study addresses the widespread lack of adherence to reporting standards such as RAT-RS in agent-based modelling (ABM) research, largely due to the time-consuming and undervalued nature of manual reporting. It presents the first systematic evaluation of the feasibility of using large language models (LLMs) to automate the generation of RAT-RS–compliant content. The authors propose a practical framework based on supervised information extraction and introduce heuristic rules to delineate model reliability and the boundaries requiring human intervention. Experimental comparisons across four LLMs demonstrate that these models produce coherent and reliable outputs for descriptive tasks, significantly enhancing reporting quality and consistency. However, they exhibit notable limitations in explanatory and evaluative tasks, underscoring the continued need for human oversight in more interpretive aspects of ABM reporting.

Adoption BarrierAgent-Based ModellingDocumentation Standards

This study addresses the widespread neglect in machine learning research of when validation occurs during data annotation—a critical factor influencing both label quality and cost—despite overreliance on post-hoc quality control. Drawing inspiration from the “shift-left” principle in software engineering, this work proposes a tripartite classification of quality checkpoints across early, intermediate, and late stages of the annotation pipeline and introduces a parameterized error propagation model that, for the first time, treats validation timing as a quantifiable design variable. Through error propagation modeling, process decomposition, and literature analysis, the authors find that only 4% of recent studies report validation timing. Their analysis demonstrates that early-stage quality checks can reduce error correction costs by up to two orders of magnitude. The paper calls for standardized reporting of timing configurations, platform support for tunable timing parameters, and empirical studies on stage-specific detection rates.

annotation pipelinesdata qualityerror propagation

Hot Scholars

CB

Conrad Borchers

Carnegie Mellon University
Educational Data MiningLearning AnalyticsIntelligent Tutoring SystemsSelf-Regulated Learning
YY

Yaxing Yao

Assistant Professor at Johns Hopkins
PrivacyIoTsHCI
PM

Philippe Muller

Professor (Computer Science), University of Toulouse, IRIT, ANITI member, TMBI
Natural Language ProcessingArtifical IntelligenceAIMachine Learning
NA

Nicholas Asher

CNRS Research Director (DRCE), member ANITI
Semanticspragmaticsdiscourse and dialogueNLP
RZ

Renwen Zhang

Assistant Professor, Nanyang Technological University
HCIMental HealthSocial SupportHealth Communication