Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls

📅 2026-07-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a key limitation in traditional web crawling analysis, which typically assumes URLs are uniformly distributed and thereby overlooks the heterogeneous “persistent core–dynamic periphery” structure inherent in real-world web graphs. To overcome this assumption, the authors propose a novel two-component urn model grounded in discovery curves and sliding windows, enabling the first quantitative estimation of both the core proportion and the dynamic evolution parameters of the periphery within a web corpus. By jointly fitting coverage and survival rates, the model demonstrates strong empirical validity on both Common Crawl and German academic web datasets. Results reveal a pronounced core–periphery organization in both collections, with further heterogeneity observed even within the dynamic periphery itself.
📝 Abstract
A longitudinal web crawl is a sequence of partial samples of an evolving URL population. Pairwise containment between two crawls is the standard probe; under a simple \emph{urn} model of the crawl -- each round samples a fraction of the URLs and replaces a fraction -- it recovers two interpretable rates, per-round survival $α$ and coverage $c$, but treats the population as uniform and consumes one pair at a time. In this work, we define a formal language for talking about a crawl. We extend this analysis with the \emph{discovery curve} $U(s, T)$, the cumulative URL footprint over a sliding window of $T$ crawls starting at $s$, which under the same urn model is also a closed-form function of $(α, c)$. Containment and the discovery curve are then two projections of one process: independent fits agree on $(α, c)$ when the urn is homogeneous, so any disagreement is itself a measurement. Applied to Common Crawl (2020--2025, domain granularity) and to the German Academic Web (GAW, URL granularity), the two projections disagree on both archives, and a two-component urn with a persistent core fraction $κ$ alongside shell parameters $(α_\partial, c_\partial)$ reconciles the disagreement. A residual on $c_\partial$ remains, signaling that the shell itself is not homogeneous; $κ$ is recorded as the scalar entry point to a rank-resolved generalization, which is left to follow-up work. \keywords{web archive \and crawl coverage \and discovery curve \and urn model \and two-component model \and URL lifetime}
Problem

Research questions and friction points this paper is trying to address.

web archive
crawl coverage
discovery curve
urn model
URL lifetime
Innovation

Methods, ideas, or system contributions that make the work stand out.

discovery curve
urn model
two-component model
crawl coverage
URL lifetime
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Michael Paris
Common Crawl Foundation, Beverly Hills, USA
H
Hande Çelikkanat
Common Crawl Foundation, Beverly Hills, USA
L
Luca Foppiano
Common Crawl Foundation, Beverly Hills, USA