Answer Generation for Questions With Multiple Information Sources in E-Commerce

📅 2021-11-27
🏛️ arXiv.org
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
In e-commerce scenarios, users pose heterogeneous, multi-source queries about products—derived from specifications, reviews, and paraphrased questions—posing challenges of information redundancy and sentiment ambiguity. Method: We propose MSQAP, an end-to-end answer generation framework featuring a novel “discriminate–fuse–generate” paradigm: (1) BERT-QA jointly models relevance and ambiguity; (2) multi-source alignment and ambiguity-aware evidence selection filters high-quality supporting evidence; and (3) T5-QA generates fluent natural-language answers. Contribution/Results: MSQAP is the first method to synergistically integrate specifications, reviews, and paraphrased questions for answer generation in e-commerce. Experiments show significant improvements: BERT-QA achieves +12.36% F1 on relevance classification; T5-QA yields +35.02% average ROUGE and +198.75% BLEU scores; and end-to-end human evaluation demonstrates +30.7% accuracy over baselines.
📝 Abstract
Automatic question answering is an important yet challenging task in E-commerce given the millions of questions posted by users about the product that they are interested in purchasing. Hence, there is a great demand for automatic answer generation systems that provide quick responses using related information about the product. There are three sources of knowledge available for answering a user posted query, they are reviews, duplicate or similar questions, and specifications. Effectively utilizing these information sources will greatly aid us in answering complex questions. However, there are two main challenges present in exploiting these sources: (i) The presence of irrelevant information and (ii) the presence of ambiguity of sentiment present in reviews and similar questions. Through this work we propose a novel pipeline (MSQAP) that utilizes the rich information present in the aforementioned sources by separately performing relevancy and ambiguity prediction before generating a response. Experimental results show that our relevancy prediction model (BERT-QA) outperforms all other variants and has an improvement of 12.36% in F1 score compared to the BERT-base baseline. Our generation model (T5-QA) outperforms the baselines in all content preservation metrics such as BLEU, ROUGE and has an average improvement of 35.02% in ROUGE and 198.75% in BLEU compared to the highest performing baseline (HSSC-q). Human evaluation of our pipeline shows us that our method has an overall improvement in accuracy of 30.7% over the generation model (T5-QA), resulting in our full pipeline-based approach (MSQAP) providing more accurate answers. To the best of our knowledge, this is the first work in the e-commerce domain that automatically generates natural language answers combining the information present in diverse sources such as specifications, similar questions, and reviews data.
Problem

Research questions and friction points this paper is trying to address.

Generating answers from multiple e-commerce data sources
Filtering irrelevant information in product-related queries
Resolving sentiment ambiguity in reviews and questions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Utilizes multiple sources: reviews, similar questions, specifications
Employs BERT-QA for relevancy prediction
Uses T5-QA for answer generation
🔎 Similar Papers
No similar papers found.
Flipkart
A
Anand A. Rajasekar
Flipkart, Bengaluru, India
N
Nikesh Garera
Flipkart, Bengaluru, India