Unapologetically Distributed: A Call for Decentralized Document Analysis

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misconception that privacy compliance constraints hinder automated document analysis and that federated learning (FL) inherently degrades performance. For the first time, we systematically demonstrate that decentralized training is not a compromise but a critical opportunity to enhance model generalization and adaptability. Methodologically, we comprehensively evaluate FL across diverse tasks, architectures, and fine-tuning strategies, while introducing multi-task fine-tuning and distributed training techniques tailored for scenarios such as table recognition. Our findings reveal that this approach significantly improves model robustness and adaptability under cross-distribution data shifts. Consequently, this work provides strong empirical support for the efficient, privacy-compliant deployment of document analysis systems, establishing decentralized learning as an enabler rather than a limitation in real-world applications.
📝 Abstract
Privacy has become an increasingly important concern in the Document Analysis community, to the extent that in many environments such as archives, governmental institutions, and local businesses, the adoption of automation is restricted by legal and policy constraints. While federated learning has often been regarded as a ``necessary evil'', implying an unavoidable performance trade-off in exchange for decentralization and privacy, many prior works overlook its potential to improve robustness to out-of-distribution data. In this paper, we present Unapologetically Distributed, the first comprehensive study evaluating distributed learning in Document Analysis along three key axes simultaneously: the tasks addressed, the architectures employed, and the fine-tuning strategies applied. Specifically, we demonstrate how various distributed training approaches enhance generalization capabilities across diverse tasks such as Table Recognition, handwriting recognition, and Word Spotting, particularly during transfer learning stages. Our results provide strong evidence that decentralization is not merely a constraint, but a valuable opportunity to improve model robustness and adaptability in real-world Document Analysis scenarios.
Problem

Research questions and friction points this paper is trying to address.

Document Analysis
Decentralized Learning
Privacy
Federated Learning
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Learning
Document Analysis
Decentralized Training
Transfer Learning
Out-of-Distribution Robustness
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Adrià Molina
Centre de Visió per Computador, Universitat Autònoma de Barcelona, Bellaterra, Catalonia; Computer Science Department, Universitat Autònoma de Barcelona, Bellaterra, Catalonia
Oriol Ramos Terrades
Oriol Ramos Terrades
Dep. Ciències de la Computació, Universitat autònoma de Barcelona - Computer Vision Centre
machine learningcomputer visionpattern recognition
Josep Lladós
Josep Lladós
Computer Vision Center, Universitat Autònoma de Barcelona
Computer VisionPattern RecognitionDocument Analysis