issue triage

Designs and operates issue-triage processes, workflows, and tools that reproduce and validate failures, generate structured bug reports with reproduction steps and metadata, classify and prioritize defects by severity and impact, assign and route issues to owners and remediation paths, and maintain benchmark issue datasets. Builds or integrates issue- and bug-tracking systems and reporting pipelines, defines defect-tracking workflows and developer metrics, and analyzes tracking data to monitor progress and coordinate remediation.

issuetriage

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.74
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Prior work lacks empirical characterization of problem-solving processes in software development. Method: Integrating grounded theory coding, sequential pattern mining, and multidimensional statistical analysis on 356 Mozilla Firefox issue reports, this study extracts fine-grained, reusable problem-solving process patterns from collaborative textual artifacts. Contribution/Results: We identify 47 empirically grounded process patterns—challenging the traditional linear assumption by revealing pervasive nonlinearity: 73% of fixes involve iterative backtracking or parallel activities. The resulting process landscape and pattern catalog systematically characterize distributional regularities across issue types, defect categories, and repair durations. This advances understanding of real-world engineering complexity and provides an evidence-based foundation for process optimization, collaborative tool design, and developer support.

Analyzing Firefox issue reportsIdentifying patterns in resolution processUnderstanding practical issue resolution

SPRINT: An Assistant for Issue Report Management

Feb 06, 2025
AA
Ahmed Adnan
🏛️ University of Dhaka | William & Mary

To address the high cost and low efficiency of manual triage for bug reports in large-scale software projects, this paper proposes a GitHub-integrated AI assistant that, for the first time, unifies multi-task deep learning across the full bug-report management pipeline—including duplicate detection, severity prediction, and fix-file recommendation. The method integrates a BERT-based semantic model, a graph neural network (GNN) for code localization, and a lightweight GitHub App architecture to deliver end-to-end interpretable recommendations. Evaluated on multiple benchmark datasets, the approach achieves F1 scores of 0.82–0.89. A real-world user study with professional developers demonstrates an average 47% reduction in triage time and a 76% recommendation adoption rate. These results significantly advance the automation level and practical utility of defect management in industrial settings.

Predicting issue severity accuratelyStreamlining issue management tasksSuggesting code modifications for issues

This work addresses the limited trust developers place in AI-generated bug reports due to their frequent lack of actionability and reproducibility. The authors propose a novel approach that integrates code coverage analysis with large language models (LLMs) to automatically detect defects in uncovered code regions and generate structured bug reports containing severity ratings, reproduction steps, and repair suggestions. A key innovation is an LLM-driven prioritization mechanism that substantially outperforms traditional rule-based methods. Evaluated on 13 Python projects, the method produced 10,467 reports; manual assessment of the top 130 revealed an 84.6% validity rate. Compared to CoverUp, it achieves higher defect validity (81.0% vs. 76.2%), a 50% improvement in P@3, and a 41% gain in mean reciprocal rank (MRR).

actionabilityAI-generated issue reportsautomated bug detection

On the Need to Monitor Continuous Integration Practices - An Empirical Study

Sep 08, 2024
JS
Jadson Santos
🏛️ Federal University of Rio Grande do Norte | University of Otago | University of Waterloo

Continuous Integration (CI) practices suffer from severe monitoring deficiencies: developers largely neglect critical metrics such as “build health” and “time-to-fix failed builds,” while mainstream CI services offer only weak native monitoring capabilities, forcing reliance on fragmented and often redundant third-party tools. Method: We conducted a triangulated investigation—including documentation analysis, developer surveys, functional audits of CI platforms, and case studies of open-source projects—to systematically identify cognitive gaps and practical monitoring needs. Contribution/Results: Our study provides the first empirical evidence that although over 80% of developers track test coverage, only a minority monitor build health or timeliness; further, all major CI services lack built-in multidimensional monitoring support. These findings establish an evidence-based foundation for designing next-generation CI monitoring frameworks and prioritizing tooling enhancements.

CI services lack native support for monitoring key practices.Developers inadequately monitor Continuous Integration practices.Third-party tools fail to fully address CI monitoring gaps.

Combining Language and App UI Analysis for the Automated Assessment of Bug Reproduction Steps

Feb 06, 2025
JM
Junayed Mahmud
🏛️ University of Central Florida | William & Mary | George Mason University

This paper addresses the challenge of low reproducibility in software bug reports caused by ambiguous or incomplete reproduction steps (S2Rs). To tackle this, we propose a cross-modal quality assessment method that synergistically integrates large language model (LLM)-based semantic understanding with dynamic UI analysis. Our approach constructs a program state model and enables fine-grained semantic alignment between S2R text descriptions and corresponding GUI interaction actions, thereby bridging the lexical diversity and programmatic semantic gap inherent in prior methods. We introduce the first LLM-driven framework that deeply unifies natural language parsing with GUI state modeling. Evaluated on standard benchmarks, our method achieves a 25.2% improvement in F1-score for S2R quality labeling and a 71.4% gain in F1-score for missing-step completion over state-of-the-art approaches—significantly enhancing both reproducibility and debugging efficiency.

Automated assessment of bug reproduction steps.Improving S2R quality with LLMs and dynamic analysis.Linking natural language to GUI interactions.

Latest Papers

What's happening recently
View more

This work addresses the challenge that bug reports in open-source projects often lack reproducible tests, hindering effective repair. The authors propose a multi-stage agent framework that decomposes test reproduction into four phases: defect localization, root cause analysis, test planning, and test generation. For the first time, this approach integrates tool-augmented mechanisms and task decomposition strategies, leveraging code-text graph retrieval, runtime environment interaction, and large language models to achieve repository-level code understanding and flexible test construction. Evaluated on SWT-bench-lite and SWT-bench-verified, the method achieves reproduction success rates of 58.43% and 70.30%, respectively—substantially outperforming existing approaches—while maintaining a low average cost of only $0.14 per instance and demonstrably enhancing downstream repair performance.

bug reproductionissue reportslarge language models

This work addresses the limitation of existing bug reports—often lacking critical information—which hinders the effectiveness of automated program repair (APR). The authors propose TrajSpec, a novel approach that introduces trajectory-guided reasoning and hierarchical evidence representation. By leveraging execution trajectories collected via proxy runs, TrajSpec extracts multi-level evidence to refine bug reports and further enhances them through integration with repository context. Combining trajectory-guided inference, hierarchical evidence modeling, and large language model–driven report generation, TrajSpec significantly improves the Pass@1 performance of multiple APR systems on SWE-Bench Lite, raising it from 47.00% to as high as 72.00%. Ablation studies confirm the contribution of each component to the overall effectiveness.

automated program repairbug reportrepair-relevant information

Hot Scholars

SC

Shing-Chi Cheung

Chair Professor of Computer Science and Engineering, HKUST
Software EngineeringSoftware TestingProgram TestingProgram Analysis
YT

Yongqiang Tian

Monash University
Software Testing and DebuggingSoftware Engineering
CF

Chunrong Fang

Software Institute, Nanjing University
Software TestingSoftware EngineeringComputer Science
MR

Manuel Rigger

National University of Singapore
Software EngineeringSystemsDatabasesProgramming Languages
RB

Ryan Beckett

Microsoft Research
networkingprogramming languagesformal methodsverification