arXiv · 2609.24092
DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents
Abstract
Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often lack direct visual counterparts on the page. Extracting these implicit values requires multi-hop derivation, such as aggregating sub-categories or reasoning over visual marks. While existing methods handle explicit text spans or simple implicit queries, they fail at multi-hop visual reasoning even after standard fine-tuning: models retrieve incorrect visual evidence, or retrieve it correctly and then skip the intermediate steps of the derivation. To address this, we introduce DocMIDE, a fine-tuning framework that trains compact vision-language models to retrieve visual evidence explicitly before deriving an answer. DocMIDE constrains generation to a plan-retrieve-derive structure and optimizes it with Group Relative Policy Optimization under a four-component, rule-based reward that scores output format, the retrieved evidence block, every intermediate derivation step, and the final value against a verified reference trace. On a 4,151-pair implicit extraction benchmark, DocMIDE raises accuracy from 70.8% to 95.9% on Qwen3.5-4B from only a small set of annotated examples, and transfers to a second backbone architecture. Supervised demonstrations alone do not close this gap at any budget we tested; rewarding the intermediate steps is what does.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jeremy Cerwin Wang, Wai Kit Wong, Jeff Kai Tai Tang. 2026-09-21. DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents. https://arxiv.org/abs/2609.24092
Cite the original work for its findings. Save a collection to share your selection of sources.