Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents

๐Ÿ“… 2026-09-17
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บไธ€็งๅˆ†่งฃ็ฝฎไฟกๅฑ‚ๆ–นๆณ•๏ผŒ้€š่ฟ‡ๆ„Ÿ็Ÿฅใ€ๅธƒๅฑ€ๅ’Œ้ชŒ่ฏไธ‰ไธช้€š้“ๆ้ซ˜้‡‘่žๆ–‡ๆกฃ็›ด้€šๅค„็†็š„ๅฏ้ ๆ€ง๏ผŒๅนถๅœจๅคšไธชๆ•ฐๆฎ้›†ไธŠ้ชŒ่ฏไบ†ๅ…ถๆœ‰ๆ•ˆๆ€งใ€‚
๐Ÿ“ Abstract
Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability for extracting the key-values, but their verbalized confidence signals are unreliable and weakly track field correctness. This paper introduces a decomposed confidence layer along three interpretable channels, including perception, layout, and validation. Together with a final conformal risk control, the score can be used for reliable STP of financial documents. The method is validated on three public datasets covering real invoices, synthetic invoices, and ad-buy forms, using two different VLM families (Qwen3.6-27B and Gemini-3.1-Flash-Lite). Our decomposed score consistently improves the separation of correct from incorrect extractions, substantially raising the AUROC from 0.54-0.74 for VLM verbalized signals to 0.90-0.99 with contributions from all three designed channels. Crucially for industrial deployment, this enables usable STP. The native VLM confidence signals could clear only 0.1%-7.0% of fields under risk control at a target error of <10%. In contrast, the proposed method auto-approves 49-72% of fields while holding the empirical error of the accepted tier at or below the target.
Problem

Research questions and friction points this paper is trying to address.

Straight-through processing
Financial documents
Calibrated probability
Residual error
Vision Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decomposed Confidence Layer
Straight-Through Processing (STP)
Vision Language Models (VLMs)
Conformal Risk Control
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Y
Yichao Jin
OCBC, Singapore
Y
Yushuo Wang
OCBC, Singapore
Yuxuan Han
Yuxuan Han
Tsinghua University
computer visioncomputer graphics
K
Kwan Ching Yee Sonia
OCBC, Singapore
W
Weiyang Song
OCBC, Singapore
C
Chiu Jin-Chun Kent
OCBC, Singapore
W
Wong Chong Hwee
OCBC, Singapore
W
Wong Tiong Kiat
OCBC, Singapore
K
Kenneth Zhu Ke
OCBC, Singapore
J
Jingyuan Zhao
OCBC, Singapore