Score
Designs and implements tools and workflows to retrieve, assemble, and restore archived web resources into coherent historical page snapshots. This includes identifying and fetching archived resources, filtering out post‑cutoff content, resolving and repairing broken links and embedded context, and synthesizing realistic page environments that reflect the original snapshot.
Web-hosted software repositories (e.g., GitHub) exhibit complex structural hierarchies that challenge long-term archiving in initiatives like the Internet Archive. Method: We crawled and analyzed over 12,000 repositories, integrating HTML integrity assessment, source tree structure parsing, and large-scale archival status comparison. Contribution/Results: This study provides the first quantitative evidence of a significant negative correlation between repository structural complexity and archival success rate. Only 4.7% of source files were archived; 31.2% of repository homepages exhibited minor rendering defects; direct-link homepage assets achieved just 14.89% archival coverage, while files in nested directories fell below 5%—revealing severe undercoverage of deeply embedded resources. Crucially, we systematically identify a structural misalignment between archival failures at the web presentation layer versus the source code layer. These findings deliver empirical grounding and concrete optimization pathways for software heritage preservation.
This study addresses the persistent loss of digital memory caused by the ephemeral nature of web content and frequent institutional website changes, exacerbated by existing archival practices that rely heavily on expert intervention and operate reactively. To counter this, the authors propose embedding proactive archiving into routine website maintenance workflows. They design and implement a lightweight, automated system leveraging Python scripts and GitHub Actions to invoke the Internet Archive’s Wayback Machine API, thereby automatically submitting web pages and associated media assets whenever site updates occur. The approach demonstrates the feasibility of low-overhead,常态化 preservation integrated into standard operations. However, the implementation also reveals that archival systems themselves remain vulnerable to platform dependencies—such as GitHub inactivity—highlighting that web ephemerality is not merely incidental but a structural condition of the contemporary web.
This study addresses a critical limitation in existing document layout analysis methods, which treat figures and tables as generic objects and thus fail to identify semantically valuable, reusable analytical visual content—referred to as “data snapshots”—in institutional documents. The work introduces the novel task of data snapshot extraction, presents a benchmark dataset comprising humanitarian reports and World Bank policy papers, and proposes an evaluation framework that integrates spatial localization with semantic annotation. Systematic evaluation of multiple open-source layout models reveals consistent shortcomings in handling institutional documents, including confusion between analytical and non-analytical content, fragmentation of composite charts, and lack of contextual awareness. By exposing the generalization bottlenecks of current models in operational documents, this research provides a foundation for future advancements through the public release of its dataset and codebase.
Existing instruction-tuning data generation methods rely heavily on high-quality seed instructions or structured web sources, limiting scalability and diversity. To address this, we propose WebR, the first framework that automatically synthesizes high-quality instruction-response pairs directly from raw, unstructured web pages—requiring no external supervision, seed data, or predefined templates. Its core innovation is a dual-perspective paradigm: “webpage-as-instruction” and “webpage-as-response,” implemented via content-driven bidirectional role modeling, unsupervised semantic alignment, and reconstruction-based generation, augmented by lightweight filtering and quality distillation. This approach significantly enhances data diversity, domain adaptability, and scalability. Evaluated on four standard instruction-following benchmarks, WebR achieves an average 16.65% improvement over prior state-of-the-art methods, demonstrating superior generalization and efficiency in domain adaptation.
Existing search agents are constrained by transient context or single-page browsing mechanisms in multi-turn complex tasks, hindering effective reuse of previously acquired information and leading to redundant retrieval and rendering. This work proposes the Fetch-then-Explore framework, which decouples page selection from evidence extraction for the first time and introduces a file system–based persistent workspace, enabling agents to revisit and extract content from previously visited pages on demand throughout the task. Built upon a unified ReAct architecture, the approach integrates three backbone models and is evaluated on the BrowseComp and WideSearch benchmarks. Experimental results demonstrate state-of-the-art accuracy on BrowseComp and consistently competitive or superior performance on WideSearch, with gains primarily attributed to efficient revisitation and secondary utilization of already accessed pages.
This study addresses the challenge of accurately estimating the absolute coverage of a web crawler over the crawlable URL space in the absence of external ground-truth data. The authors propose a statistical method that relies solely on longitudinal crawl data from a single crawler. By analyzing the intersections of URLs across multiple consecutive crawls, they formulate an urn-model-based estimation framework and employ linear regression to infer the coverage ratio. Notably, the approach requires neither external benchmarks nor comparisons across multiple crawlers, making it applicable to any focused longitudinal crawling scenario. Experiments on 15 semi-annual crawls of the German academic web from 2013 to 2021 demonstrate that, under stable configurations, the crawler achieves approximately 46% coverage, thereby validating the method’s effectiveness and practical utility.
This work addresses the challenges of manual GitHub dataset construction, where evidence and annotations are scattered across multiple sources, hindering traceability and collaborative auditing. The authors propose the first end-to-end framework that integrates browser-assisted data collection, local structured storage, and decision provenance tracking. Leveraging a Chrome extension to capture page snapshots, comments, and labels, the system couples an Express backend with a React dashboard to consolidate all research decisions into a unified SQLite workspace, tightly linking evidence with explanatory context. Evaluated on the Matplotlib project, the framework successfully collected 22 snapshots, 38 comments, and 98 annotations across 20 issues, fully preserving the research process and enabling conflict simulation. This approach substantially enhances the auditability and reproducibility of dataset construction.