🤖 AI Summary
This work addresses the challenges of fine-grained product information extraction from shopping review webpages and the difficulty of dynamically updating product databases. To this end, we propose MarkupLM++, the first model to extend sequence labeling to internal nodes—not only leaf nodes—of the DOM tree, thereby explicitly modeling hierarchical webpage structure and semantic associations. Built upon the MarkupLM architecture, MarkupLM++ incorporates large-scale, diverse web annotation data and explicit DOM tree structure for end-to-end fine-tuning. Experiments on a real-world e-commerce review dataset yield 90.6% precision, 72.4% recall, and 80.5% F1-score, significantly outperforming baseline methods. This work advances web document understanding from “textual content extraction” toward “structured semantic parsing,” establishing a new paradigm for constructing real-time-updating product knowledge bases and enabling downstream applications such as customer analytics and recommendation.
📝 Abstract
Web information extraction (WIE) is an important part of many e-commerce systems, supporting tasks like customer analysis and product recommendation. In this work, we look at the problem of building up-to-date product databases by extracting detailed information from shopping review websites. We fine-tuned MarkupLM on product data gathered from review sites of different sizes and then developed a variant we call MarkupLM++, which extends predictions to internal nodes of the DOM tree. Our experiments show that using larger and more diverse training sets improves extraction accuracy overall. We also find that including internal nodes helps with some product attributes, although it leads to a slight drop in overall performance. The final model reached a precision of 0.906, recall of 0.724, and an F1 score of 0.805.