arXiv · 2610.11402
GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA
Abstract
GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness. We present \textbf{CoVeR-VQA}, a training-free multi-stage verification and correction framework for grounded multi-view VQA. Starting from GPT-5.6 zero-shot predictions, CoVeR-VQA progressively applies view-specific verification with Gemini-3.6-Flash, prior-guided joint verification with Claude-Opus-5, and cross-split group-level verification that exploits semantically filtered question groups from shared multi-view scenes and validation-derived prior knowledge. On the official GoldenViewVQA test set, the four-stage CoVeR-VQA pipeline achieves 84.75\% Joint Accuracy, improving the GPT-5.6 zero-shot baseline by 13.56 percentage points, while reaching 94.92\% Answer Accuracy and 86.44\% View Accuracy. The final submitted run achieves 88.14\% Joint Accuracy after two additional evaluator-informed post-hoc corrections. Our analysis shows that supporting-view localization remains the primary source of residual errors, highlighting the importance of explicit evidence verification for reliable multi-view multimodal reasoning.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kun Wang, Yupeng Hu, Ruping Cao, Hao Liu, Zhiran Li, Qianlong Xiang, Harry Cheng. 2026-10-08. GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA. https://arxiv.org/abs/2610.11402
Cite the original work for its findings. Save a collection to share your selection of sources.