Empirical Analysis and Detection of Hallucinations in LLM-Generated Bug Report Summaries

📅 2026-05-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) are prone to hallucinations—such as omitting critical information or fabricating details—when generating structured summaries of software bug reports, potentially misleading developers and undermining tool reliability. This work presents the first systematic analysis of such hallucinations from a section-aware perspective. The authors introduce a controllable synthetic hallucination injection mechanism to construct a benchmark dataset and propose a multi-task joint detection framework that simultaneously predicts whether a report contains hallucinations, localizes the affected sections, and identifies the hallucination type. Experiments on the BugsRepo dataset demonstrate that the best-performing model achieves Macro-F1 scores of 0.89, 0.83, and 0.84 at the report, section, and hallucination-type levels, respectively, effectively uncovering prevalent hallucination patterns and underlying model failure mechanisms.
📝 Abstract
Large Language Models (LLMs) are increasingly used to generate summaries of software bug reports, including sections such as Steps-to-Reproduce (S2R), Actual Behavior (AB), and Expected Behavior (EB). However, these models frequently produce hallucinations that can be convincing but unsupported by the source report. This can mislead developers and reduce trust in automated maintenance tools. Existing hallucination detection approaches typically evaluate outputs at the full-response level and do not consider the structure of technical documents. An initial exploratory study on 80 structured bug report summaries found that approximately 47.9% contained missing information, while 12.3% included fabricated content, highlighting the need for systematic hallucination analysis in bug report summarization. In this work, we empirically investigate hallucinations in LLM-generated bug report summaries from a section-aware perspective. Using the BugsRepo dataset, derived from Mozilla OSS projects, we introduce controlled synthetic hallucination injection to construct a benchmark for training and evaluation. We propose a section-aware hallucination detection approach that jointly predicts whether a summary contains hallucinated content, identifies affected sections, and classifies hallucination types. Experimental results across multiple pretrained language models show that the proposed approach achieves strong performance across all tasks, with the best model obtaining 0.89 report-level Macro-F1, 0.83 section-level Macro-F1, and 0.84 hallucination-type Macro-F1. We further analyze common hallucination patterns and model failure modes to better understand limitations of current LLM-generated bug report summaries. The findings highlight the importance of section-aware hallucination analysis for improving the reliability of LLM-assisted bug report summarization in software maintenance workflows.
Problem

Research questions and friction points this paper is trying to address.

hallucination
bug report summarization
large language models
software maintenance
structured technical documents
Innovation

Methods, ideas, or system contributions that make the work stand out.

section-aware hallucination detection
structured bug report summarization
synthetic hallucination injection
hallucination type classification
empirical analysis of LLM hallucinations
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Hinduja Nirujan
Electrical and Software Engineering, University of Calgary, Calgary, Alberta, Canada
S
Shreyas Patil
Electrical and Software Engineering, University of Calgary, Calgary, Alberta, Canada
A
Abdallah Ayoub
Electrical and Software Engineering, University of Calgary, Calgary, Alberta, Canada
A
Ahmad Abdel Latif
Electrical and Software Engineering, University of Calgary, Calgary, Alberta, Canada
G
Gouri Ginde
Electrical and Software Engineering, University of Calgary, Calgary, Alberta, Canada