arXiv · 2609.32036
ScreenHaystack: Finding Blind Zones in GUI Grounding
Abstract
We introduce ScreenHaystack, a dynamic needle-in-a-haystack benchmark for evaluating spatial reliability in GUI grounding. Instead of testing each target at a fixed position, ScreenHaystack systematically relocates controlled target icons across high-resolution GUI backgrounds and measures whether models can localize them consistently. Using this benchmark, we find that leading GUI grounding models, including Qwen3-VL, UI-TARS, GTA, and UI-Venus, exhibit blind zones: spatial regions where grounding accuracy drops sharply despite fixed target appearance and instruction. These blind zones transfer to unseen ScreenSpot-Pro examples: targets inside blind zones are consistently harder to ground, with Qwen3-VL-8B dropping by 16.1 percentage points, and controlled relocation shows that moving targets into blind zones decreases accuracy while moving them out improves accuracy. We further show through controlled synthetic experiments that uneven spatial coverage in training data can induce such blind zones. Therefore, we propose a simple strategy, blind-zone-oriented augmentation, which adds supervision in blind zones and improves ScreenSpot-Pro accuracy over both original and randomly augmented Click-100k fine-tuning.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chenyue Li, Xiaoxiao Sun, Yubo Deng, Qinlin Zhao, Serena Yeung-Levy, Yuhui Zhang. 2026-09-25. ScreenHaystack: Finding Blind Zones in GUI Grounding. https://arxiv.org/abs/2609.32036
Cite the original work for its findings. Save a collection to share your selection of sources.