Learning Page Order in Shuffled WOO Releases

📅 2026-02-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of automatically reordering pages in Dutch Freedom of Information (WOO) PDF documents, which suffer from shuffled page sequences and heterogeneous content—including emails, legal texts, and tables—rendering semantic cues unreliable. The authors propose a page-embedding-based approach and systematically compare Pointer Networks, seq2seq Transformers, and a specialized pairwise ranking model. Their findings reveal that short and long documents require fundamentally different reordering strategies. By tailoring models to document length, they achieve a substantial performance gain for long documents (Kendall’s τ improves by 0.21), whereas generic seq2seq models severely degrade on longer inputs (τ drops to 0.014). The best-performing method attains τ scores ranging from 0.95 for 2–5-page documents to 0.72 for 15-page documents, while also uncovering why curriculum learning fails in this context.

Technology Category

Machine Learning: Learning Preferences or RankingsNatural Language Processing: Learning & Optimization for NLPSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
We investigate document page ordering on 5,461 shuffled WOO documents (Dutch freedom of information releases) using page embeddings. These documents are heterogeneous collections such as emails, legal texts, and spreadsheets compiled into single PDFs, where semantic ordering signals are unreliable. We compare five methods, including pointer networks, seq2seq transformers, and specialized pairwise ranking models. The best performing approach successfully reorders documents up to 15 pages, with Kendall's tau ranging from 0.95 for short documents (2-5 pages) to 0.72 for 15 page documents. We observe two unexpected failures: seq2seq transformers fail to generalize on long documents (Kendall's tau drops from 0.918 on 2-5 pages to 0.014 on 21-25 pages), and curriculum learning underperforms direct training by 39% on long documents. Ablation studies suggest learned positional encodings are one contributing factor to seq2seq failure, though the degradation persists across all encoding variants, indicating multiple interacting causes. Attention pattern analysis reveals that short and long documents require fundamentally different ordering strategies, explaining why curriculum learning fails. Model specialization achieves substantial improvements on longer documents (+0.21 tau).
Problem

Research questions and friction points this paper is trying to address.

document page ordering
shuffled documents
WOO releases
page reordering
heterogeneous documents
Innovation

Methods, ideas, or system contributions that make the work stand out.

document page ordering
seq2seq transformers
curriculum learning failure
learned positional encodings
attention pattern analysis
🔎 Similar Papers
No similar papers found.