issue mining

Design and build data-extraction and processing pipelines that collect, parse, link, and deduplicate issue reports, comments, events, labels, and related artifacts from issue trackers and across repositories. Produce cleaned, structured datasets and metadata exports suitable for querying, analysis, or integration with downstream tools.

issuemining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Jul 30, 2025
MD
Mattia Di Profio
🏛️ University of Aberdeen

Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.

Automating ETL workflows to reduce human interventionDesigning context-specific transformations without manual inputStandardizing diverse datasets using example-driven approaches

This work addresses the challenge of reconciling high throughput and low query latency in traditional ETL pipelines when processing continuously arriving fresh data, where unpredictable preprocessing operations often create bottlenecks. The authors propose Fluid ETL Pipelines, which introduce, for the first time, an elastic and non-blocking preprocessing mechanism that decouples data ingestion from transformation. By dynamically scheduling preprocessing tasks based on resource availability and user interest—without blocking data ingestion—and leveraging preemptible computing resources such as Amazon Spot instances, the approach significantly reduces operational costs. Experimental results demonstrate that Fluid ETL Pipelines substantially improve the efficiency of exploring fresh data, offering a novel direction for accelerating real-time queries and enabling adaptive preprocessing management.

data preprocessing routinesETL pipelinesfresh data exploration

This study addresses the lack of systematic understanding regarding the challenges users encounter in developing and maintaining nf-core standardized bioinformatics pipelines. Conducting the first large-scale empirical analysis, we examined 25,173 GitHub issues and pull requests using BERTopic for topic modeling, Cohen’s δ effect size statistics, and data mining techniques, identifying 13 key problem categories spanning core challenges such as tool development, CI configuration, and containerization debugging. Our findings reveal that 89.38% of reported issues were ultimately resolved, with half addressed within three days. Furthermore, the presence of issue labels and code snippets significantly enhanced resolution efficiency, offering empirical evidence to inform strategies for improving the sustainability and collaborative effectiveness of nf-core pipelines.

GitHub issuesnf-corepipeline maintenance

DataLens: ML-Oriented Interactive Tabular Data Quality Dashboard

Jan 28, 2025
MA
Mohamed Abdelaal
🏛️ Software AG | TU Darmstadt

Existing data management tools suffer from limited automation, poor interactivity, and insufficient integration with ML workflows, compromising data quality and hindering analytical and modeling performance. To address this, we propose an interactive, ML-oriented tabular data quality dashboard that establishes an adaptive, human-in-the-loop + ML-driven data cleaning闭环. Our approach integrates data profiling, multi-strategy error detection and repair—including statistical analysis, rule-based engines, and supervised/semi-supervised models—while supporting expert rule validation and labeling. Cleaning strategies are iteratively refined using downstream model performance feedback. Furthermore, we unify DataSheets, MLflow, and Delta Lake to ensure reproducibility, traceability, and versioning of the cleaning pipeline. Experiments across multiple benchmark datasets demonstrate significant improvements: error identification rate and repair accuracy increase notably, downstream ML models achieve an average 7.2% accuracy gain, and cleaning time decreases by 40%.

Data ManagementData QualityMachine Learning Integration

Data Processing for the OpenGPT-X Model Family

Oct 11, 2024
NB
Nicolo’ Brandizzi
🏛️ Fraunhofer IAIS | Fraunhofer IIS | DFKI

This work addresses data quality, multilingual coverage, and regulatory compliance challenges in training large language models (LLMs) for the OpenGPT-X initiative. Methodologically, it introduces a novel “dual-track” data processing paradigm: lightweight filtering for curated datasets and aggressive filtering combined with MinHash/LSH-based deduplication for large-scale web corpora—fully aligned with EU regulations such as the GDPR. The pipeline integrates fastText-based language identification, hybrid rule-and-statistics filtering, a learned quality scoring model, and end-to-end metadata provenance tracking. Its primary contribution is the construction of the first high-quality, EU-compliant multilingual corpus for LLM training, explicitly designed for public-sector applications. Empirical evaluation demonstrates substantial improvements in model robustness, transparency, and auditability—particularly in government and public service use cases—while ensuring legal and ethical adherence across 24 official EU languages.

Develop data pipeline for multilingual OpenGPT-X LLMsEnsure compliance with European data regulationsHandle curated and web data with distinct processing methods

Latest Papers

What's happening recently
View more

Community-driven scientific workflow ecosystems often struggle to sustain themselves due to ambiguous maintenance and user support mechanisms, particularly in cross-platform collaboration and heterogeneous execution environments. This study presents the first cross-platform empirical analysis of the nf-core ecosystem, systematically examining 15,760 GitHub issues, 35,411 pull requests, and 895 forum discussions. By integrating metadata and textual features into predictive models, the research uncovers significant disparities in maintenance and support activities across platforms and highlights weak explicit linkages among them. The findings reveal that issues, pull requests, and forum posts predominantly serve distinct roles—coordinating maintenance, facilitating code integration, and providing user support, respectively. Moreover, issue actionability, diagnostic evidence, and depth of interaction emerge as critical determinants of resolution efficiency.

community-drivenheterogeneous execution environmentsmaintenance

This work addresses the critical impact of unresolved issue artifacts—such as bugs and missing documentation—on software quality in open-source projects. To overcome the limitations of existing tools, which lack systematic analysis of issue lifecycles and evolutionary patterns, we propose G-Issue, the first mining tool that integrates issue lifecycle modeling with evolution tracking. Built on a Python API, G-Issue efficiently collects and analyzes issue data from open-source repositories, achieving faster mining performance than mainstream tools while uncovering key patterns in issue evolution. The tool further enables issue prioritization and developer assignment based on evolutionary characteristics, offering both a novel methodology and empirical evidence to support quality management in open-source software development.

issue evolutionissue lifetimeissue-related artifacts

This work addresses the challenge of efficiently accessing structured scholarly publications and associated software metadata within research knowledge bases. To this end, the authors propose a generic and extensible interoperable pipeline architecture built upon the shared Grid’5000/ABACA infrastructure, integrating modules for document preprocessing, information extraction, software mention recognition, and visualization. Designed to support multi-team collaboration, user validation, and external interoperability, the system demonstrates its utility through a daily tracking application of software mentions in the HAL open archive, significantly enhancing the visibility of research software and advancing open science practices. Experimental results confirm that the pipeline efficiently processes large-scale scientific literature and enables automated extraction and visual representation of software mentions within the HAL portal.

open scienceresearch repositoriesscientific literature

Current data processing pipelines for post-training large language models—encompassing cleaning, deduplication, synthesis, and quality filtering—are fragmented and lack auditability and sample-level decision transparency. This work proposes the first end-to-end configurable data processing framework that unifies data ingestion, cleaning, LLM-driven synthesis across eight task types, three-tiered quality gating, and export modules. The system introduces sample-level provenance tracking and a precise hallucination verification mechanism. It supports six input formats and over 100 model APIs via LiteLLM, offering both a YAML-driven command-line interface and a Python API. Outputs are compatible with five training formats used by TRL, Unsloth, and AlignTune, substantially enhancing transparency, reproducibility, and scalability in post-training data preparation.

data curationLLM post-trainingpipeline auditability

Existing large language model (LLM)-driven data analysis tools are often confined to isolated subtasks and struggle to support end-to-end executable analytical workflows. This work proposes an autonomous, sandboxed, and auditable end-to-end system that leverages LLMs for action planning, iteratively generating structured operations, executing code in a secure environment, and integrating streaming traceability with intermediate result previews. By unifying a structured action backend, sandboxed execution, and an interactive visual interface—features integrated here for the first time—the system enables users to drive complete analytical workflows using only natural language. Users can inspect, modify, and export the entire process and its outputs directly within a web browser, ensuring full reproducibility, editability, and transparency throughout the analytical pipeline.

action tracedata analysisend-to-end workflow

Hot Scholars

BA

Bram Adams

Queen's University
software release engineeringsoftware integrationsoftware build systemssoftware modularity
RG

Raula Gaikovina Kula

Professor, The University of Osaka
Software EcosystemsDeveloper ProficiencySoftware in SocietySoftware Engineering
MS

Mojtaba Shahin

Assistant Professor in Software Engineering, RMIT University
AI EngineeringEmpirical Software EngineeringSoftware ArchitectureDevOps
PL

Peng Liang

School of Computer Science, Wuhan University
Software EngineeringSoftware ArchitectureEmpirical Software Engineering
YZ

Yanjie Zhao

Huazhong University of Science and Technology
Software EngineeringSoftware Security