Semantic Navigation for Issue Localization in Code Repository

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficient iterative fault localization in code repositories by LLM-based agents, which stems from a lack of structured navigation and evidential support. To this end, it proposes SemNav, a framework that integrates deterministic retrieval with large language model reasoning. Specifically, SemNav constructs a semantic navigation graph that parses dependencies on demand to ensure comprehensive candidate coverage, generates problem-oriented compact semantic cards to facilitate precise fault correction, and establishes a persistent workspace grounded in Language Server Protocol (LSP) static analysis to support dynamic evidence verification. Experimental results demonstrate that SemNav achieves a File Hit@10 of 82.67%, reduces context overhead by 48.2%, and improves the downstream issue resolution rate to 52.33%, comprehensively outperforming existing baselines.
📝 Abstract
Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and revise their judgments about these candidates as new evidence is acquired. Existing environments, however, provide limited support for this loop: agents must search for unresolved relation targets, reconstruct entity semantics from raw source code, and revise candidates without evidential basis. To address these limitations, we present SemNav, a framework that leverages deterministic retrieval to seed a broad candidate set and an LLM agent to continually refine that set, thereby combining initial coverage with evidence-guided revision. SemNav supports this process through three key components. A Semantic Navigation Graph resolves program relations on demand through a language server, enabling direct navigation to related entities across files. Issue-conditioned Semantic Cards provide compact, source-grounded interpretations of each entity's role and relevance to the issue. A persistent Candidate Workspace records each candidate together with its evidential basis, enabling grounded verification, revision, and ranking. Across SWE-bench Lite and PLocBench, SemNav outperforms existing baselines, improving File Hit@10 from 68.33\% to 82.67\% with Gemma 4B. Component ablations and trajectory analysis support the complementary roles of all three components, while Semantic Cards reduce working-context load by 48.2\% relative to full-source reading. SemNav further ranks first on all seven evidence-quality metrics on SWE-Explore and improves downstream issue resolution from 44.00\% to 52.33\%.
Problem

Research questions and friction points this paper is trying to address.

issue localization
semantic navigation
code repository
LLM agents
evidence-guided revision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Navigation Graph
Issue Localization
LLM Agent
Semantic Cards
Candidate Workspace
🔎 Similar Papers
No similar papers found.