Robust Hierarchical Structures for Agentic Document Analysis

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency and lack of extraction guarantees in LLM agents caused by neglecting document hierarchical structure. To overcome these limitations, this work proposes SHED, a two-stage workflow grounded in theoretical document space representations and tool-calling agents. By leveraging a robust and compact hierarchy inference algorithm, SHED optimizes the document analysis pipeline and provides the first theoretical guarantees for structure extraction, effectively balancing robustness with compactness. Experimental results demonstrate that SHED reduces computational costs by 90% while improving F1 scores by 13%–68% and accuracy by 3%–23%, significantly enhancing task-relevant filtering capabilities.
📝 Abstract
Large Language Models (LLMs) enable us to better understand text documents, including PDFs and Word documents. However, LLMs, as well as more modern LLM agents, i.e., those with tool-calling abilities, typically treat such documents as plain text, ignoring the fact that they are often organized hierarchically into sections and subsections. Extracting this structure, while difficult, can improve efficiency and effectiveness for agents (and humans)---since only sections relevant to a given task need to be processed. Unfortunately, prior work on structure extraction provides no formal guarantees on how well the inferred structure matches the true one. Instead, we target a robust and compact variant that is feasible to infer and useful in practice. Robustness ensures that the text under each subsection header is a superset of the text under the same header in the true structure. Compactness seeks to minimize this superset, reducing agentic cost (or human cognitive load). We propose SHED, a two-stage workflow for inferring a robust and compact structure. The first stage is pluggable with an infinite family of approaches, each guaranteeing robustness for a specific document class. We theoretically characterize the document space using these classes and their hierarchical relationships. Empirically, SHED improves F-1 scores (measuring the robustness--compactness trade-off) by 13%--68% over non-LLM baselines and 9%--15% over expensive LLM-based approaches. Finally, we show how SHED-inferred structures are valuable for agentic document analysis: agents using SHED outperform baselines, achieving 3%--23% higher accuracy while being up to 10x cheaper.
Problem

Research questions and friction points this paper is trying to address.

Document Structure Extraction
Hierarchical Document Analysis
Large Language Models
LLM Agents
Robustness and Compactness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Structure Extraction
Robustness and Compactness
Agentic Document Analysis
Two-stage Workflow
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruiying Ma
UC Berkeley
Yiming Lin
Yiming Lin
UC Berkeley
Data ManagementData QualityUnstructured Data ExplorationQuery Optimization
A
Aditya G. Parameswaran
UC Berkeley