arXiv · 2609.21351
PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR
Abstract
Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell content alignment. A Bayesian decision strategy maximizes alignment accuracy by exploiting the correlation between extraction errors and computable rule violations. Evaluated on open-source and custom VLMs, PrismAlign reduces hallucinations and achieves state-of-the-art performance on OmniDocBench 1.5, as well as on the table category of CC-OCR and PureDocBench.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Guangyi Liu, Qianjun Huang, Boyu Hou. 2026-09-18. PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR. https://arxiv.org/abs/2609.21351
Cite the original work for its findings. Save a collection to share your selection of sources.