WCXB: A Multi-Type Web Content Extraction Benchmark

📅 2026-05-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing benchmarks for main content extraction from web pages, which are typically small-scale, homogeneous, and outdated, thereby failing to adequately evaluate system generalization across diverse page structures. To overcome this, the authors introduce WCXB, a new benchmark comprising 2,008 pages from 1,613 domains spanning seven structurally distinct categories, including news articles, forums, and product pages. High-quality annotations are ensured through a rigorous five-stage pipeline combining LLM-assisted labeling, automated validation, four rounds of model-based review, and human verification. Evaluation of 13 state-of-the-art extraction systems reveals that while top-performing methods achieve an F1 score of 0.93 on article pages, their performance drops substantially on non-news structured pages (F1 = 0.41–0.84), highlighting a critical generalization gap. The dataset is publicly released.
📝 Abstract
Web content extraction - isolating a page's main content from surrounding boilerplate - is a prerequisite for search indexing, retrieval-augmented generation, NLP dataset construction, and large language model training. Progress in this area has been constrained by the limitations of existing evaluation benchmarks, which are small (100-800 pages), restricted to news articles, or based on web pages from over a decade ago. We introduce the Web Content Extraction Benchmark (WCXB), a dataset of 2,008 web pages from 1,613 domains spanning seven structurally distinct page types: articles, forums, products, collections, listings, documentation, and service pages. The dataset includes a 1,497-page development set and a 511-page held-out test set with matched page type distributions. Ground truth annotations were produced through a five-stage pipeline: LLM-assisted drafting, automated verification, four-pass frontier model review, snippet and quality verification scripts, and human review. We evaluate 13 extraction systems - 11 heuristic and 2 neural - and find that while top systems converge on articles (F1 = 0.93), performance diverges sharply on structured page types (F1 = 0.41-0.84), revealing blind spots invisible to existing article-only benchmarks. The dataset is released under CC-BY-4.0 with HTML source files, ground truth annotations, page type labels, and baseline results.
Problem

Research questions and friction points this paper is trying to address.

web content extraction
evaluation benchmark
multi-type web pages
boilerplate removal
dataset limitation
Innovation

Methods, ideas, or system contributions that make the work stand out.

web content extraction
benchmark dataset
multi-type web pages
LLM-assisted annotation
structured page evaluation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Murrough Foley