arXiv Science⌕ Search

arXiv · 2610.11553

EVIE: Evidence-Vector-Informed Embeddings for Visual Document Retrieval

Abstract

Accurate and scalable visual document retrieval (VDR) requires both fine-grained page understanding and efficient indexing, yet existing approaches struggle to achieve both. OCR-based text retrieval adds preprocessing latency and can lose visual and structural cues needed to understand complex pages. Single-vector vision-language models bypass OCR, but compressing an entire page into one vector limits the granularity of query--document matching. Multi-vector retrievers with MaxSim provide finer interactions, yet demand large indexes and still leave room for accuracy improvements. We argue that overcoming these limitations requires preserving query-relevant page evidence throughout representation learning and index construction. To this end, we introduce \textbf{\textit{EVIE}} (Evidence-Vector-Informed Embeddings), a family of native visual document retrievers integrating three key innovations: (1) Evidence-judged data governance, which uses a multimodal judge to identify answer-bearing positives and filter unreliable negatives. (2) Bidirectional teacher--student learning with symmetric listwise distillation and prefix-based Matryoshka representation learning (Prefix-MRL), enabling one student checkpoint to serve six nested embedding dimensions without re-encoding. (3) Hierarchical agglomerative index compression (HAC), which clusters page tokens with spatial regularization and stores semantic centroids for single-stage MaxSim retrieval. Extensive experiments across 138 tasks from ViDoRe V1, V2, V3, and JinaVDR validate EVIE. EVIE-8B achieves 66.75 nDCG@10 on V3, exceeding the best external baseline by 1.43 points, with a four-suite average of 79.51. EVIE-4.5B with HAC retains 59.58 nDCG@10 at only 3.81 GiB per million pages, reducing vector payload by $128\times$. Together, these results improve the accuracy--storage trade-off for visual document retrieval.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zifei Wang, Wei Wen, Qiang Ji, Qian-Wen Zhang, Ruizhi Qiao, Xing Sun. 2026-10-08. EVIE: Evidence-Vector-Informed Embeddings for Visual Document Retrieval. https://arxiv.org/abs/2610.11553

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

LIME: Link-based User-item Interaction Modeling with Decoupled XOR Attention for Efficient Test Time Scaling

Scaling large recommendation systems requires advancing three major frontiers: processing longer user histories, expanding candidate sets, and increasing model capacity. While promising, transformers' computational cost scales quadratically with the user sequence length and linearly with the number of candidates. This trade-off makes it prohibitively expensive to expand candidate sets or increase sequence length at inference, despite the significant performance improvements. We introduce \textbf{LIME}, a novel architecture that resolves this trade-off. Through two key innovations, LIME fundamentally reduces computational complexity. First, low-rank ``link embeddings" enable pre-computation of attention weights by decoupling user and candidate interactions, making the inference cost nearly independent of candidate set size. Second, a linear attention mechanism, \textbf{LIME-XOR}, reduces the complexity with respect to user sequence length from quadratic ($O(N^2)$) to linear ($O(N)$). Experiments on public and industrial datasets show LIME achieves near-parity with state-of-the-art transformers but with a 10$\times$ inference speedup on large candidate sets or long sequence lengths. When tested on a major recommendation platform, LIME improved user engagement while maintaining minimal inference costs with respect to candidate set size and user history length, establishing a new paradigm for efficient and expressive recommendation systems.

cs.IR↗

Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval

A skill can match a task's topic while conflicting with its resource, procedure, or output requirements. We study this as same-capability risk-exposure retrieval and introduce SameCapRisk-Bench: 890 units and 1,314 query cases across five mechanisms and twelve conflict types. Each unit pairs a skill that meets a query requirement with a same-capability skill that violates it, with evidence for that distinction. The evaluation tests source-task requirements and two kinds of paired queries that reverse which skill fits: changing the requested evidence role or exact output interface while holding the skills fixed. To capture both retrieval success and conflicting exposure, Recall tracks helpful hits, harmful sibling rate (HSR) tracks exposure of the conflicting sibling, and CleanHit requires a helpful hit without that exposure. Four public skill retrievers expose conflicting siblings at HSR@3 of 0.737-0.881 on source-task contracts, compared with 0.099-0.12 on controlled source-role queries. Across the fixed mixture of 1,235 held-out queries, their Recall@3 is 0.903-0.944 and HSR@3 is 0.344-0.393. The conflicting sibling ranks first on 24.8-29.2% of source-task queries for these four retrievers. We also examine what the reranker receives: truncating long skills can remove the passages where the two skills' contracts differ. With the same BGE top-20 candidates, increasing the reranker's input budget from 512 to 4,096 tokens reduces source-task sibling-first errors by 8.27 pp, with an uncertain CleanHit gain.

cs.IR↗

RecToolBench: Benchmarking Recommendation-Specific Tool Orchestration under Fuzzy User Intent

Recent advances in agentic recommender systems are shifting recommender systems from passive filtering engines to instruction-following agents that use external tools to resolve user intent. However, existing benchmarks often assume explicit user intent, simplified tool environments, or isolated function calls, leaving realistic tool orchestration for recommendation underexplored. To bridge this gap, we propose RecToolBench, a Model Context Protocol (MCP)-based benchmark for evaluating tool-using recommender agents under fuzzy user instructions. RecToolBench contains more than 1,200 executable tasks across three recommendation domains, 13 MCP servers, and 32 tools, spanning single-tool calls, parallel tool calls, sequential tool chains, and hybrid tool orchestration. We construct RecToolBench with a scalable synthesize--fuzzify--judge pipeline that generates executable fuzzy recommendation tasks, and evaluates agent trajectories using rule-based execution checks and rubric-based LLM evaluation. Experiments on representative LLMs show that syntactically valid tool calls do not guarantee successful recommendations. Models struggle with semantic parameter grounding, multi-step evidence integration, and grounded final recommendations, especially as orchestration complexity increases. Our results identify tool orchestration under fuzzy user intent as a major bottleneck for agentic recommender systems. Our data and code are available at https://github.com/ShawnChenn/RecToolBench.

cs.IR↗