Markup Language Modeling for Web Document Understanding

📅 2025-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of fine-grained product information extraction from shopping review webpages and the difficulty of dynamically updating product databases. To this end, we propose MarkupLM++, the first model to extend sequence labeling to internal nodes—not only leaf nodes—of the DOM tree, thereby explicitly modeling hierarchical webpage structure and semantic associations. Built upon the MarkupLM architecture, MarkupLM++ incorporates large-scale, diverse web annotation data and explicit DOM tree structure for end-to-end fine-tuning. Experiments on a real-world e-commerce review dataset yield 90.6% precision, 72.4% recall, and 80.5% F1-score, significantly outperforming baseline methods. This work advances web document understanding from “textual content extraction” toward “structured semantic parsing,” establishing a new paradigm for constructing real-time-updating product knowledge bases and enabling downstream applications such as customer analytics and recommendation.

Technology Category

Natural Language Processing: Information ExtractionData Mining & Knowledge Management: WebMachine Learning: Structured Learning

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Web information extraction (WIE) is an important part of many e-commerce systems, supporting tasks like customer analysis and product recommendation. In this work, we look at the problem of building up-to-date product databases by extracting detailed information from shopping review websites. We fine-tuned MarkupLM on product data gathered from review sites of different sizes and then developed a variant we call MarkupLM++, which extends predictions to internal nodes of the DOM tree. Our experiments show that using larger and more diverse training sets improves extraction accuracy overall. We also find that including internal nodes helps with some product attributes, although it leads to a slight drop in overall performance. The final model reached a precision of 0.906, recall of 0.724, and an F1 score of 0.805.
Problem

Research questions and friction points this paper is trying to address.

Extracting product information from shopping review websites
Building up-to-date product databases for e-commerce systems
Improving web information extraction accuracy using DOM nodes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuned MarkupLM model on product data
Extended predictions to DOM tree internal nodes
Used larger diverse training sets for accuracy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Su Liu
B
Bin Bi
J
Jan Bakus
P
Paritosh Kumar Velalam
V
Vijay Yella
V
Vinod Hegde