arXiv · 2610.02880
Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models
Abstract
Retrieval-augmented document question answering assumes that once the right page is found, a vision-language model (VLM) can read it. We show that this assumption often fails, leaving a retrieval-reading gap: evidence found but not used. A paired protocol isolates this gap by comparing answers from the retrieved page images alone with answers from the same images plus their extracted text. On FoveDoc-Bench, our benchmark with traceable evidence, retrieval finds nearly every evidence page, yet adding CPU-OCR text raises strict accuracy by 13 to 16 points. An exact text layer roughly doubles the gain, which appears across six VLMs from three families and, within one family, narrows with scale without closing. The reader can read this evidence but cannot find it: crops of it recover most of the text gain, boxes around it on the page none. The same protocol identifies two boundaries. Extracted text helps on textual evidence but is neutral or harmful on charts and figures. Its advantage shrinks as retrieval degrades, and unrelated text of the same form adds nothing detectable. Extracted text is an amplifier of retrieval that works, not a substitute for retrieval that does not. Our code is available at https://github.com/atoz03/fovedoc-sup.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Qingtao Xia, Siyao Cheng, Jiahua Bao, Jiaxing Du, Jie Liu. 2026-10-02. Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models. https://arxiv.org/abs/2610.02880
Cite the original work for its findings. Save a collection to share your selection of sources.