PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决表格OCR中的结构错误和语义幻觉问题,提出PrismAlign方法,通过多视角视觉语言模型对齐并结合表格逻辑先验来提高准确性。
📝 Abstract
Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell content alignment. A Bayesian decision strategy maximizes alignment accuracy by exploiting the correlation between extraction errors and computable rule violations. Evaluated on open-source and custom VLMs, PrismAlign reduces hallucinations and achieves state-of-the-art performance on OmniDocBench 1.5, as well as on the table category of CC-OCR and PureDocBench.
Problem

Research questions and friction points this paper is trying to address.

table extraction
structural errors
semantic hallucinations
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-VLM framework
prior-steered alignment
Bayesian decision strategy
hallucination-robust
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Guangyi Liu
Huawei Technologies, Co., Ltd
Q
Qianjun Huang
Huawei Technologies, Co., Ltd
B
Boyu Hou
Huawei Technologies, Co., Ltd