Score
Design and implement methods and pipelines that rank and prioritize program functions or code regions by importance or risk (hotspot ranking, vulnerability prioritization) so downstream tools or reviewers focus on the highest-impact locations. These systems compute and combine repo-wide signals to direct automated or human analysis, with the goal of reducing expensive analysis calls and lowering developer triage effort.
Facing the dual challenges of rapidly increasing vulnerability volumes and constrained resources, existing vulnerability prioritization methods lack both a unified theoretical framework and practical deployability. This study conducts a systematic literature review (SLR) of 82 primary works to establish, for the first time, a five-dimensional unified taxonomy encompassing severity, exploitability, contextual relevance, predictive capability, and aggregation mechanisms. The analysis reveals critical bottlenecks: insufficient cross-domain generalizability, weak adaptability to dynamic environments, and low industrial integration. It further identifies dynamism, context awareness, and scalability as core future research directions. To bridge the structural gap between academic research and real-world practice, the study proposes a reusable evaluation framework and a comprehensive technology roadmap. These contributions advance both theoretical understanding and operational applicability in vulnerability management.
Software issue triage faces challenges of low operational efficiency and a persistent academic–industrial gap in complex system maintenance. To address this, we conduct a systematic literature review (SLR) spanning 234 English and Chinese publications from 2004 to 2023. This is the first SLR to jointly analyze academic research and industrial practice, revealing three critical bottlenecks: misaligned objectives between academia and industry, lack of standardized evaluation criteria, and practical deployment barriers. We propose a unified triage evaluation framework structured along four dimensions—data, tasks, metrics, and benchmarks—and systematically catalog open-source datasets and empirical methodologies to enable reproducible performance validation. All reviewed literature and supporting resources are publicly released. Our work establishes a foundational theory, provides actionable evaluation tools, and outlines collaborative pathways to bridge the gap between laboratory research and industrial adoption of triage technologies.
Software source code often harbours"hotspots": small portions of the code that change far more often than the rest of the project and thus concentrate maintenance activity. We mine the complete version histories of 91 evolving, actively developed GitHub repositories and identify 15 recurring line-level hotspot patterns that explain why these hotspots emerge. The three most prevalent patterns are Pinned Version Bump (26%), revealing brittle release practices; Long Line Change (17%), signalling deficient layout; and Formatting Ping-Pong (9%), indicating missing or inconsistent style automation. Surprisingly, automated accounts generate 74% of all hotspot edits, suggesting that bot activity is a dominant but largely avoidable source of noise in change histories. By mapping each pattern to concrete refactoring guidelines and continuous integration checks, our taxonomy equips practitioners with actionable steps to curb hotspots and systematically improve software quality in terms of configurability, stability, and changeability.
Implementation discrepancies across software repository mining tools severely threaten the validity of empirical findings. Method: We conduct a dual-tool comparative analysis of 10 large-scale open-source projects, systematically identifying how minor implementation differences—such as commit parsing logic and author deduplication rules—induce up to 500% deviation in key metrics (e.g., commit count, developer count). We propose a “tool-level configuration + post-hoc normalization” co-optimization framework to mitigate metric divergence and perform multi-tool experiments, quantitative consistency assessment, and code-level root-cause analysis. Contribution/Results: We identify six technical sources undermining data validity and establish the first validity assessment paradigm for Mining Software Projects Research (MSPR) explicitly addressing tool heterogeneity—thereby enabling rigorous, reproducible, and comparable empirical software engineering studies.
In large monorepos, CI merge pipelines under high load become integration bottlenecks, severely impeding development velocity. To address this, we propose a build-system-agnostic optimization framework: leveraging historical build logs, PR metadata, and contextual features, we train a lightweight model to predict per-PR build success probability; based on these predictions, we dynamically prioritize pull requests—scheduling high-probability requests first during peak loads. Our approach requires no modifications to underlying CI infrastructure, ensuring strong integrability and deployment feasibility. Evaluated on a real-world large-scale production monorepo, it achieves significantly higher throughput compared to FIFO and non-learning baselines, effectively alleviating CI integration bottlenecks. The method provides a practical, scalable optimization pathway for high-concurrency software delivery without compromising system compatibility or operational simplicity.
This work addresses the intractability of program analysis, testing, and optimization in modern languages like Rust due to combinatorial explosion in configuration spaces. To tackle this challenge, the authors propose a compiler-based approach for prioritizing configurations. By instrumenting the Rust compiler to extract intermediate representations, they construct a configuration dependency graph and refine configuration rankings using graph centrality metrics combined with code impact scope. A SAT solver is then employed to generate a high-relevance subset of configurations while preserving their semantic validity. The prototype system, RustyEx, efficiently produces valid configuration subsets of specified sizes on mainstream Rust projects, significantly improving the exploration efficiency of large configuration spaces under resource constraints.
This work addresses the inefficiency of current AI programming systems that uniformly employ high-end large language models for all tasks, incurring excessive inference costs for routine software engineering activities. The authors propose a dynamic routing mechanism that leverages code health metrics and task metadata to assign tasks to the lowest-cost model tier—lightweight, standard, or heavyweight—that meets required quality thresholds. Introducing code quality signals into model selection for the first time, they formulate verifiable cost-effectiveness conditions and establish a rigorous evaluation protocol to quantify the impact of individual code health subfactors on routing decisions. Evaluated on SWE-bench Lite using a three-tier strategy—comprising heuristic thresholds, machine learning classifiers, and an ideal oracle—the approach demonstrates significant cost reduction without compromising output quality when lightweight models achieve pass rates on high-health code exceeding their relative cost advantage and the code health effect size reaches 𝑝̂ ≥ 0.56.
This study addresses the lack of large-scale empirical analysis on the impact of hotfixes in real-world development contexts, particularly amid the rise of autonomous coding agents. It introduces, for the first time, a repository-level operational definition of urgency and conducts a systematic investigation of hotfix practices across more than 61,000 GitHub repositories by integrating code change mining with comparative analyses of human and AI-driven behaviors. The findings reveal that hotfixes are typically characterized by single-author contributions, minimal code modifications, low engagement in code review, and rarely involve test changes. Moreover, the study uncovers over ten statistically significant behavioral differences between human developers and AI agents in hotfix scenarios, offering crucial empirical insights for future human-AI collaborative software maintenance.
This work addresses the scarcity of fine-grained, expert-annotated non-functional requirement (NFR) samples in existing public GitHub datasets by introducing GitReq—the first large-scale GitHub-based quality requirements dataset aligned with the ISO/IEC 25010 standard. GitReq comprises 6,302 expert-validated requirements extracted from 4,080 repositories, spanning all eight quality characteristics defined in the standard. The construction methodology employs category-specific tri-signal mining, a preprocessing step to separate functional from non-functional requirements, and a rigorous manual annotation protocol, achieving a Fleiss’ Kappa inter-annotator agreement of 0.72. Experimental evaluation demonstrates that GPT-5.2 attains a macro-averaged F1 score of 0.641 under zero-shot settings, confirming both the dataset’s validity and its inherent challenge for current language models.
This work addresses the inefficiency of existing deployment freeze policies, which fail to differentiate change risk during live events or rapid releases, and the limitations of conventional risk prediction approaches that rely on developer metadata or extensive historical data—raising privacy concerns and suffering from poor generalizability. To overcome these challenges, the authors propose a diff-aware risk assessment framework that extracts both quantitative and qualitative features, such as structural complexity, directly from code changes. Notably, it leverages large language models (LLMs) as cross-language feature extractors for risk prediction, eliminating dependence on language-specific tooling and preserving developer privacy. Empirical evaluation demonstrates the approach’s effectiveness, achieving an average recall of 0.83 and an F1-score of 0.81 on both Prime Video’s production environment and the ApacheJIT dataset, confirming its robustness across multi-language and multi-organizational settings.