Score
Designs and builds question-answering systems that identify and describe differences between two or more states, versions, or views and return direct answers to user queries about what changed. Also develops capabilities to explain reasons or attributes of those changes, support multiple question types, and enable fine-grained, interactive interrogation of changes.
This work addresses the challenge of reduced question-answering accuracy in multi-version software systems, where documentation across versions is highly similar yet contains subtle differences that confuse existing QA systems. To tackle this issue, the authors propose QAMR, a novel chatbot that introduces a retrieval-augmented generation (RAG) framework specifically tailored for multi-version documentation. The framework incorporates a dual-chunking strategy—optimizing chunks separately for retrieval and generation—along with query rewriting and context selection mechanisms. Evaluated on both real-world industrial data and public benchmarks, QAMR achieves a question-answering accuracy of 88.5% and a retrieval accuracy of 90%, representing improvements of 16.5% and 12% over baseline methods, respectively, while also reducing response time by 8%.
This study presents the first systematic investigation of revision behaviors for Architecture-Related Questions (ARQs) on Stack Overflow, aiming to understand how users dynamically refine architecturally significant information. Based on a manually annotated dataset of 4,114 ARQs with revision histories, we employ qualitative coding, empirical analysis, and statistical testing. We find that ARQ revisions are infrequent yet highly time-sensitive—most occur before the first answer and are predominantly performed by askers. Fourteen categories of supplementary information are identified, with design context, component dependencies, and architectural concerns being the most prevalent. Primary revision motivations center on “clarifying architectural understanding” and “improving question readability.” Empirical evaluation demonstrates that revisions significantly enhance answer practicality, informativeness, and relevance. Our findings provide novel evidence for community-driven mechanisms underlying architectural knowledge evolution in Q&A platforms.
This study addresses the persistent challenge in multi-turn question answering where user intent is ambiguous, particularly the difficulty in generating accurate answers even after clarification. The work decouples this problem into two phases: clarification strategy and post-clarification answering. It proposes a supervised fine-tuning approach to optimize clarification strategies and introduces the PACIFIC evaluation framework to systematically analyze performance across both phases. Findings reveal that while clarification strategies can be effectively improved through fine-tuning, models still exhibit substantially low answer accuracy after correctly executing clarifications, highlighting that accurately interpreting users’ clarification responses constitutes a critical bottleneck in achieving intent alignment.
Existing conversational systems suffer from low diversity and insufficient informativeness in follow-up question generation, leading to ambiguous user interactions. Method: This paper proposes a novel approach that leverages large language models (LLMs) to generate “hypothetical comprehensive answers” as an intermediate representation, explicitly modeling the information gap between the user’s current knowledge state and the ideal answer—thereby enabling targeted generation of highly informative and coverage-rich follow-up questions. Contribution/Results: To our knowledge, this is the first work to integrate counterfactual answer generation, explicit information-gap modeling, and supervised fine-tuning for cognitively grounded question prompting. On multiple benchmarks, the fine-tuned model achieves a 32% improvement in question diversity and a 27% increase in information entropy; human evaluation confirms statistically significant superiority over state-of-the-art baselines, validating both the effectiveness and novelty of explicit information-gap modeling for enhancing question quality.
In e-commerce scenarios, users pose heterogeneous, multi-source queries about products—derived from specifications, reviews, and paraphrased questions—posing challenges of information redundancy and sentiment ambiguity. Method: We propose MSQAP, an end-to-end answer generation framework featuring a novel “discriminate–fuse–generate” paradigm: (1) BERT-QA jointly models relevance and ambiguity; (2) multi-source alignment and ambiguity-aware evidence selection filters high-quality supporting evidence; and (3) T5-QA generates fluent natural-language answers. Contribution/Results: MSQAP is the first method to synergistically integrate specifications, reviews, and paraphrased questions for answer generation in e-commerce. Experiments show significant improvements: BERT-QA achieves +12.36% F1 on relevance classification; T5-QA yields +35.02% average ROUGE and +198.75% BLEU scores; and end-to-end human evaluation demonstrates +30.7% accuracy over baselines.
This work addresses the lack of a unified analytical framework for retrieval and reasoning pipelines in multi-hop question answering, which has hindered systematic comparison across methods. We propose, for the first time, a four-axis design framework that treats the execution process as the fundamental unit of analysis, encompassing execution plans, index structures, control strategies, and termination criteria. Through a comprehensive literature review and ablation studies, we structurally map prominent approaches—including RAG and agent-based systems—onto this framework using benchmarks such as HotpotQA, revealing consistent trade-offs among effectiveness, efficiency, and evidence faithfulness. Our framework systematically organizes existing design choices, identifies reproducible empirical trends, and highlights key challenges, including structure-aware planning and transferable control strategies.
This work addresses the challenge of outdated documentation and fragmented design knowledge in rapidly evolving codebases, where critical insights are scattered across source code and pull requests. To tackle this, the authors propose a method for incrementally constructing a typed software knowledge graph grounded in pull request data. Leveraging schema-driven semantic extraction and relation inference, the system supports three retrieval modalities—including agent-guided graph traversal—enabling question-answering over a living documentation system and tracking semantic change histories. This approach represents the first effort to automatically transform pull requests into a structured, queryable, and continuously evolving knowledge base. Evaluation on a real-world Ruby on Rails project demonstrates the system’s ability to generate highly relevant, evidence-backed answers, though user feedback indicates room for improvement in the conciseness of synthesized documentation.
This study addresses the lack of systematic representation and comparison of rhetorical strategies employed by humans and large language models (LLMs) in retrieval-based question answering. It proposes DiscoTrace, a novel method grounded in Rhetorical Structure Theory that models answers as sequences of rhetorical acts paired with question interpretations, thereby establishing the first structured framework enabling direct comparison of response strategies across human communities and LLMs. Integrating discourse act annotation, question interpretation modeling, and cross-group analysis, the approach reveals significant variation in rhetorical strategies among nine distinct human communities. In contrast, even under imitation prompting, LLMs exhibit limited rhetorical diversity and consistently favor broad coverage over selective engagement with specific question interpretations.