Score
Designs, implements, and evaluates end-to-end question-answering systems and their components, including answer-generation and response-generation modules, interactive dialogue or QA flows, and system architectures for retrieving or generating answers. Also builds and curates question banks and questionnaires, automates question/answer workflows, and defines and runs answer-quality evaluation and benchmarking pipelines.
Traditional question-answering (QA) systems suffer from strong data dependency and poor adaptability to dynamic environments. Method: This paper proposes a large language model (LLM)-based agent architecture for QA, introducing the first hierarchical LLM agent framework specifically designed for QA tasks. It comprises four stages—planning, question understanding, information retrieval, and answer generation—and integrates multi-step task decomposition, external tool invocation, and retrieval-augmented generation (RAG). Contribution/Results: The framework significantly enhances interactive reasoning capabilities and cross-environment generalization compared to conventional pipeline-based and naive LLM-based QA approaches. Through systematic evaluation, we identify key performance bottlenecks and distill three critical future research directions: scalability, trustworthy reasoning, and environment-aware collaboration. This work provides both theoretical foundations and a structured roadmap for advancing LLM-agent-driven QA systems.
To address data scarcity, fragmented workflows, and insufficient evaluation in domain-adaptive question answering (QA) model development, this paper introduces the first integrated QA data generation–fine-tuning–evaluation closed-loop platform. Leveraging large language models (LLMs), the platform enables context-aware, adaptive QA pair synthesis; supports interactive dataset browsing and model exploration; and provides multi-dimensional evaluation metric visualization alongside cross-model performance benchmarking. Its key contributions are: (1) the first end-to-end, auditable closed-loop framework for domain QA; (2) tight integration of data quality assessment with model behavior analysis; and (3) support for local deployment and full workflow reproducibility. Experimental results demonstrate that the platform significantly improves both development efficiency and interpretability of domain-specific QA models. The source code will be publicly released.
Formal email reply generation is time-consuming, imposes high cognitive load, and critically depends on users’ ability to craft effective prompts. Method: This paper proposes a question-answering (QA)-driven large language model (LLM) interaction paradigm: the system automatically parses incoming emails to extract semantic intent and generates concise, structured questions; users answer these questions, and the system synthesizes a contextually appropriate, professional reply. Contribution/Results: This work pioneers the decomposition of high-level text generation into accessible, low-threshold QA interactions—eliminating the need for manual prompt engineering. Controlled experiments and field studies demonstrate that the approach significantly improves response efficiency and substantially reduces cognitive load, while maintaining parity with conventional prompt-based methods in politeness, content completeness, and domain-specific professionalism.
In e-commerce scenarios, users pose heterogeneous, multi-source queries about products—derived from specifications, reviews, and paraphrased questions—posing challenges of information redundancy and sentiment ambiguity. Method: We propose MSQAP, an end-to-end answer generation framework featuring a novel “discriminate–fuse–generate” paradigm: (1) BERT-QA jointly models relevance and ambiguity; (2) multi-source alignment and ambiguity-aware evidence selection filters high-quality supporting evidence; and (3) T5-QA generates fluent natural-language answers. Contribution/Results: MSQAP is the first method to synergistically integrate specifications, reviews, and paraphrased questions for answer generation in e-commerce. Experiments show significant improvements: BERT-QA achieves +12.36% F1 on relevance classification; T5-QA yields +35.02% average ROUGE and +198.75% BLEU scores; and end-to-end human evaluation demonstrates +30.7% accuracy over baselines.
This study investigates whether AI-generated educational assessments can match human-authored items in psychometric quality and user satisfaction. Method: We developed an automated NLP course quiz generation pipeline using GPT-4o-mini and conducted the first integrated psychometric evaluation combining unidimensional and multidimensional Item Response Theory (IRT) with Differential Item Functioning (DIF) analysis. A mixed-methods assessment was performed from both student and domain-expert perspectives. Contribution/Results: LLM-generated items demonstrated strong discrimination and appropriate difficulty levels, with stable IRT parameter estimates. DIF analysis identified only two potentially biased items requiring review. Students and experts rated item quality, clarity, and pedagogical alignment highly (mean ≥4.3/5). This work establishes a methodological framework and empirical foundation for AI-powered, scalable, interpretable, and psychometrically sound educational assessment.
This work addresses critical limitations of large language models (LLMs) in complex question answering (CQA)—including low accuracy, poor interpretability, uncontrolled knowledge integration, and frequent hallucinations—particularly in high-stakes domains such as multi-objective energy policy decision-making. To overcome these challenges, we propose the first systematic hybrid CQA architecture, integrating five synergistic techniques: domain adaptation, multi-step task decomposition, neuro-symbolic fusion, human-in-the-loop reinforcement supervision, and program synthesis. Our framework employs structured knowledge anchoring, iterative decomposition, and multimodal retrieval augmentation to enhance cross-cultural reasoning and multi-objective decision-making. We further introduce a rigorous evaluation benchmark emphasizing fairness, robustness, and anti-hallucination capabilities. The study establishes a paradigm shift toward trustworthy, auditable, and human-intervenable CQA systems driven by controllable hybrid architectures, providing both theoretical foundations and practical guidelines for next-generation intelligent QA.
This work addresses the challenge of reduced question-answering accuracy in multi-version software systems, where documentation across versions is highly similar yet contains subtle differences that confuse existing QA systems. To tackle this issue, the authors propose QAMR, a novel chatbot that introduces a retrieval-augmented generation (RAG) framework specifically tailored for multi-version documentation. The framework incorporates a dual-chunking strategy—optimizing chunks separately for retrieval and generation—along with query rewriting and context selection mechanisms. Evaluated on both real-world industrial data and public benchmarks, QAMR achieves a question-answering accuracy of 88.5% and a retrieval accuracy of 90%, representing improvements of 16.5% and 12% over baseline methods, respectively, while also reducing response time by 8%.
Overreliance on large language models (LLMs) for pedagogical question generation in learning analytics suffers from inefficiency, opacity, and misalignment with instructional objectives. Method: This paper proposes a two-stage “generate-verify” framework leveraging small language models (SLMs). In the first stage, an SLM generates diverse candidate questions; in the second, probabilistic verification and re-ranking—guided by structured reasoning—select questions exhibiting high answer definiteness and strong pedagogical alignment. Contribution/Results: To our knowledge, this is the first work to deeply integrate SLM-based text generation with probabilistic inference for educational question generation, thereby extending SLMs’ capabilities in complex instructional tasks. Evaluated via dual human–machine assessment (seven domain experts + LLM-based evaluation), the method achieves LLM-level performance in answer clarity and learning objective consistency, demonstrating that lightweight models—when embedded in a carefully designed architecture—can deliver high-fidelity, educationally grounded question generation.
Existing QA system testing methods suffer from two key limitations: (1) synthetically generated questions lack naturalness and fail to trigger real-world defects, and (2) reliance on static datasets restricts question diversity and contextual relevance. To address these, we propose CQ²A, a context-driven question generation framework that innovatively integrates large language models (LLMs) with semantic context modeling. CQ²A first extracts entities and relations from input contexts to construct realistic answers, then prompts an LLM to generate natural, contextually grounded test questions. It further incorporates consistency verification and constraint checking to ensure high-quality output. Extensive experiments across three benchmark datasets demonstrate that CQ²A significantly improves defect detection rate, question naturalness, and context coverage. Moreover, fine-tuning QA systems with CQ²A-generated test cases substantially reduces error rates, validating its practical utility in robustness evaluation and model improvement.
Existing open-domain question answering systems struggle to support users in iteratively refining and deeply exploring initial answers, primarily due to the absence of mechanisms that generate relevant insights to enrich the interactive experience. This work introduces, for the first time, a document-level insight generation task tailored to open-ended questions, accompanied by the SCOpE-QA dataset. The authors propose InsightGen, a two-stage framework that first constructs a document topic graph via clustering and then selects contextual neighborhoods from this graph to prompt large language models to produce diverse, relevant, and actionable supplementary insights. Experimental results across 3,000 questions demonstrate that the approach effectively extends or reconstructs initial answers, establishing a strong baseline for this novel task.
This work addresses the limitation of existing approaches in multi-hop question generation, which overlook the intrinsic duality between question generation and question answering, thereby constraining generation quality. To overcome this, the paper proposes the QQ framework, which explicitly models this dual relationship for the first time by jointly training multi-hop question generation and question answering within a unified architecture. The framework incorporates bidirectional alignment constraints and a contrastive learning mechanism to strengthen semantic correspondence between generated questions and their answers. Experimental results demonstrate that the proposed method significantly improves question quality on the HotpotQA and MuSiQue datasets, with both automatic metrics and human evaluations confirming its superiority over baseline approaches.