The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the accuracy–energy trade-off arising from image versus text input modalities when processing privacy-sensitive documents with local small models. Leveraging models of ≤8B parameters and FP8 quantization, this work systematically evaluates the energy efficiency of vision-language and text-only models across diverse document types. It proposes a strategy to dynamically switch input modalities based on document layout complexity, thereby circumventing energy-intensive neural OCR, and identifies batch processing as a primary energy-saving mechanism. Experimental results demonstrate that batching reduces energy consumption by 38%–85% without compromising accuracy. Ultimately, this research establishes document-type-specific best practice guidelines for efficient, privacy-compliant information extraction.
📝 Abstract
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input representation, model family, and inference configuration. Benchmarking on the near-plain-text Kleister-NDA contracts and the layout-rich VRDU forms, we find that batching is the dominant energy lever, cutting energy per page by 38-85% at no cost in accuracy, while FP8 quantization saves 27-32% when requests are served one at a time but less than 1mWh per page (9-19%) once batching is applied. Preprocessing dominates what remains: neural OCR costs $17\times$ more energy per page than classical OCR and never reaches the Pareto frontier. Which representation wins flips with the type of document: vision--language models on layout-rich documents and small text-only models with a cheap parser on near-plain text, where they are both more accurate and cheaper than any vision--language configuration. Our work yields concrete guidelines for energy-efficient, privacy-compliant local information extraction.
Problem

Research questions and friction points this paper is trying to address.

Information Extraction
Energy Efficiency
Privacy-sensitive Documents
Small Language Models
Accuracy-Energy Trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Information Extraction
Energy Efficiency
Vision-Language Models
Small Local Models
Accuracy-Energy Trade-off
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.