🤖 AI Summary
Large language models (LLMs) are prone to hallucinations—such as omitting critical information or fabricating details—when generating structured summaries of software bug reports, potentially misleading developers and undermining tool reliability. This work presents the first systematic analysis of such hallucinations from a section-aware perspective. The authors introduce a controllable synthetic hallucination injection mechanism to construct a benchmark dataset and propose a multi-task joint detection framework that simultaneously predicts whether a report contains hallucinations, localizes the affected sections, and identifies the hallucination type. Experiments on the BugsRepo dataset demonstrate that the best-performing model achieves Macro-F1 scores of 0.89, 0.83, and 0.84 at the report, section, and hallucination-type levels, respectively, effectively uncovering prevalent hallucination patterns and underlying model failure mechanisms.
📝 Abstract
Large Language Models (LLMs) are increasingly used to generate summaries of software bug reports, including sections such as Steps-to-Reproduce (S2R), Actual Behavior (AB), and Expected Behavior (EB). However, these models frequently produce hallucinations that can be convincing but unsupported by the source report. This can mislead developers and reduce trust in automated maintenance tools. Existing hallucination detection approaches typically evaluate outputs at the full-response level and do not consider the structure of technical documents. An initial exploratory study on 80 structured bug report summaries found that approximately 47.9% contained missing information, while 12.3% included fabricated content, highlighting the need for systematic hallucination analysis in bug report summarization. In this work, we empirically investigate hallucinations in LLM-generated bug report summaries from a section-aware perspective. Using the BugsRepo dataset, derived from Mozilla OSS projects, we introduce controlled synthetic hallucination injection to construct a benchmark for training and evaluation. We propose a section-aware hallucination detection approach that jointly predicts whether a summary contains hallucinated content, identifies affected sections, and classifies hallucination types. Experimental results across multiple pretrained language models show that the proposed approach achieves strong performance across all tasks, with the best model obtaining 0.89 report-level Macro-F1, 0.83 section-level Macro-F1, and 0.84 hallucination-type Macro-F1. We further analyze common hallucination patterns and model failure modes to better understand limitations of current LLM-generated bug report summaries. The findings highlight the importance of section-aware hallucination analysis for improving the reliability of LLM-assisted bug report summarization in software maintenance workflows.