build question-answering systems

Designs, implements, and evaluates end-to-end question-answering systems and their components, including answer-generation and response-generation modules, interactive dialogue or QA flows, and system architectures for retrieving or generating answers. Also builds and curates question banks and questionnaires, automates question/answer workflows, and defines and runs answer-quality evaluation and benchmarking pipelines.

buildquestion-answeringsystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.73
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

QGen Studio: An Adaptive Question-Answer Generation, Training and Evaluation Platform

Apr 08, 2025
MM
Movina Moses
🏛️ IBM Research | University College London

To address data scarcity, fragmented workflows, and insufficient evaluation in domain-adaptive question answering (QA) model development, this paper introduces the first integrated QA data generation–fine-tuning–evaluation closed-loop platform. Leveraging large language models (LLMs), the platform enables context-aware, adaptive QA pair synthesis; supports interactive dataset browsing and model exploration; and provides multi-dimensional evaluation metric visualization alongside cross-model performance benchmarking. Its key contributions are: (1) the first end-to-end, auditable closed-loop framework for domain QA; (2) tight integration of data quality assessment with model behavior analysis; and (3) support for local deployment and full workflow reproducibility. Experimental results demonstrate that the platform significantly improves both development efficiency and interpretability of domain-specific QA models. The source code will be publicly released.

Enables custom QA dataset creation using LLMsFacilitates model fine-tuning on synthetic dataSupports model comparison and performance benchmarking

Understanding and Supporting Formal Email Exchange by Answering AI-Generated Questions

Feb 06, 2025
YM
Yusuke Miura
🏛️ Waseda University | The University of Tokyo

Formal email reply generation is time-consuming, imposes high cognitive load, and critically depends on users’ ability to craft effective prompts. Method: This paper proposes a question-answering (QA)-driven large language model (LLM) interaction paradigm: the system automatically parses incoming emails to extract semantic intent and generates concise, structured questions; users answer these questions, and the system synthesizes a contextually appropriate, professional reply. Contribution/Results: This work pioneers the decomposition of high-level text generation into accessible, low-threshold QA interactions—eliminating the need for manual prompt engineering. Controlled experiments and field studies demonstrate that the approach significantly improves response efficiency and substantially reduces cognitive load, while maintaining parity with conventional prompt-based methods in politeness, content completeness, and domain-specific professionalism.

Enhances email reply efficiencyMaintains email qualityReduces cognitive workload

In e-commerce scenarios, users pose heterogeneous, multi-source queries about products—derived from specifications, reviews, and paraphrased questions—posing challenges of information redundancy and sentiment ambiguity. Method: We propose MSQAP, an end-to-end answer generation framework featuring a novel “discriminate–fuse–generate” paradigm: (1) BERT-QA jointly models relevance and ambiguity; (2) multi-source alignment and ambiguity-aware evidence selection filters high-quality supporting evidence; and (3) T5-QA generates fluent natural-language answers. Contribution/Results: MSQAP is the first method to synergistically integrate specifications, reviews, and paraphrased questions for answer generation in e-commerce. Experiments show significant improvements: BERT-QA achieves +12.36% F1 on relevance classification; T5-QA yields +35.02% average ROUGE and +198.75% BLEU scores; and end-to-end human evaluation demonstrates +30.7% accuracy over baselines.

Filtering irrelevant information in product-related queriesGenerating answers from multiple e-commerce data sourcesResolving sentiment ambiguity in reviews and questions

Evaluating LLM-Generated Q&A Test: a Student-Centered Study

May 10, 2025
AW
Anna Wróblewska
🏛️ Warsaw University of Technology | MODUL University Vienna

This study investigates whether AI-generated educational assessments can match human-authored items in psychometric quality and user satisfaction. Method: We developed an automated NLP course quiz generation pipeline using GPT-4o-mini and conducted the first integrated psychometric evaluation combining unidimensional and multidimensional Item Response Theory (IRT) with Differential Item Functioning (DIF) analysis. A mixed-methods assessment was performed from both student and domain-expert perspectives. Contribution/Results: LLM-generated items demonstrated strong discrimination and appropriate difficulty levels, with stable IRT parameter estimates. DIF analysis identified only two potentially biased items requiring review. Students and experts rated item quality, clarity, and pedagogical alignment highly (mean ≥4.3/5). This work establishes a methodological framework and empirical foundation for AI-powered, scalable, interpretable, and psychometrically sound educational assessment.

Compares LLM-generated and human-authored test performanceDevelops AI pipeline for reliable Q&A test generationEvaluates GPT-4o-mini test quality with students and experts

Complex QA and language models hybrid architectures, Survey

Feb 17, 2023
XD
Xavier Daull
🏛️ Naval Group | Toulon Université | Aix Marseille Univ | CNRS | LIS

This work addresses critical limitations of large language models (LLMs) in complex question answering (CQA)—including low accuracy, poor interpretability, uncontrolled knowledge integration, and frequent hallucinations—particularly in high-stakes domains such as multi-objective energy policy decision-making. To overcome these challenges, we propose the first systematic hybrid CQA architecture, integrating five synergistic techniques: domain adaptation, multi-step task decomposition, neuro-symbolic fusion, human-in-the-loop reinforcement supervision, and program synthesis. Our framework employs structured knowledge anchoring, iterative decomposition, and multimodal retrieval augmentation to enhance cross-cultural reasoning and multi-objective decision-making. We further introduce a rigorous evaluation benchmark emphasizing fairness, robustness, and anti-hallucination capabilities. The study establishes a paradigm shift toward trustworthy, auditable, and human-intervenable CQA systems driven by controllable hybrid architectures, providing both theoretical foundations and practical guidelines for next-generation intelligent QA.

Addressing specialized requirements like domain knowledge and reasoningOvercoming LLM limitations for complex question-answering tasksReviewing hybrid architectures and strategies for improved QA performance

Latest Papers

What's happening recently
View more

This work addresses the challenge of reduced question-answering accuracy in multi-version software systems, where documentation across versions is highly similar yet contains subtle differences that confuse existing QA systems. To tackle this issue, the authors propose QAMR, a novel chatbot that introduces a retrieval-augmented generation (RAG) framework specifically tailored for multi-version documentation. The framework incorporates a dual-chunking strategy—optimizing chunks separately for retrieval and generation—along with query rewriting and context selection mechanisms. Evaluated on both real-world industrial data and public benchmarks, QAMR achieves a question-answering accuracy of 88.5% and a retrieval accuracy of 90%, representing improvements of 16.5% and 12% over baseline methods, respectively, while also reducing response time by 8%.

multi-release systemsquestion answeringretrieval-augmented generation

Overreliance on large language models (LLMs) for pedagogical question generation in learning analytics suffers from inefficiency, opacity, and misalignment with instructional objectives. Method: This paper proposes a two-stage “generate-verify” framework leveraging small language models (SLMs). In the first stage, an SLM generates diverse candidate questions; in the second, probabilistic verification and re-ranking—guided by structured reasoning—select questions exhibiting high answer definiteness and strong pedagogical alignment. Contribution/Results: To our knowledge, this is the first work to deeply integrate SLM-based text generation with probabilistic inference for educational question generation, thereby extending SLMs’ capabilities in complex instructional tasks. Evaluated via dual human–machine assessment (seven domain experts + LLM-based evaluation), the method achieves LLM-level performance in answer clarity and learning objective consistency, demonstrating that lightweight models—when embedded in a carefully designed architecture—can deliver high-fidelity, educationally grounded question generation.

Assess question quality via human experts and large language modelsDevelop a question generation pipeline using small language modelsGenerate and validate questions through probabilistic reasoning

Testing Question Answering Software with Context-Driven Question Generation

Nov 11, 2025
SL
Shuang Liu
🏛️ Renmin University of China | Tianjin University

Existing QA system testing methods suffer from two key limitations: (1) synthetically generated questions lack naturalness and fail to trigger real-world defects, and (2) reliance on static datasets restricts question diversity and contextual relevance. To address these, we propose CQ²A, a context-driven question generation framework that innovatively integrates large language models (LLMs) with semantic context modeling. CQ²A first extracts entities and relations from input contexts to construct realistic answers, then prompts an LLM to generate natural, contextually grounded test questions. It further incorporates consistency verification and constraint checking to ensure high-quality output. Extensive experiments across three benchmark datasets demonstrate that CQ²A significantly improves defect detection rate, question naturalness, and context coverage. Moreover, fine-tuning QA systems with CQ²A-generated test cases substantially reduces error rates, validating its practical utility in robustness evaluation and model improvement.

Existing methods ignore context limiting question diversityGenerating unnatural questions reduces bug detection effectivenessTesting approaches lack reliability in real-world scenario coverage

Existing open-domain question answering systems struggle to support users in iteratively refining and deeply exploring initial answers, primarily due to the absence of mechanisms that generate relevant insights to enrich the interactive experience. This work introduces, for the first time, a document-level insight generation task tailored to open-ended questions, accompanied by the SCOpE-QA dataset. The authors propose InsightGen, a two-stage framework that first constructs a document topic graph via clustering and then selects contextual neighborhoods from this graph to prompt large language models to produce diverse, relevant, and actionable supplementary insights. Experimental results across 3,000 questions demonstrate that the approach effectively extends or reconstructs initial answers, establishing a strong baseline for this novel task.

answer refinementdocument-grounded QAopen-ended QA

This work addresses the limitation of existing approaches in multi-hop question generation, which overlook the intrinsic duality between question generation and question answering, thereby constraining generation quality. To overcome this, the paper proposes the QQ framework, which explicitly models this dual relationship for the first time by jointly training multi-hop question generation and question answering within a unified architecture. The framework incorporates bidirectional alignment constraints and a contrastive learning mechanism to strengthen semantic correspondence between generated questions and their answers. Experimental results demonstrate that the proposed method significantly improves question quality on the HotpotQA and MuSiQue datasets, with both automatic metrics and human evaluations confirming its superiority over baseline approaches.

intrinsic dualitymulti-hop question generationnatural language generation

Hot Scholars

MB

Mohit Bansal

Parker Distinguished Professor, Computer Science, UNC Chapel Hill
Natural Language ProcessingComputer VisionMachine LearningMultimodal AI
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
DM

Dinesh Manocha

Distinguished University Professor, University of Maryland at College Park
computer graphicsgeometric modelingmotion planningvirtual reality
YH

Yangfan He

University of Minnesota - Twin Cities
AI AgentReasoningAI AlignmentFoundation Models
YH

Yao Hu

浙江大学
Machine Learning