arXiv · 2605.27800
CuriosAI Submission to the CASTLE Challenge at EgoVis 2026
Abstract
CASTLE 2026 asks 185 multiple-choice questions over 600+ hours of synchronized multi-view egocentric video. We explore two approaches on top of a shared multimodal preprocessing layer, including per-person timelines, speaker-resolved transcripts, and multi-VLM caption ensembles. Approach A, SVA: Search-Verify-Answer, is a three-stage pipeline that hierarchically narrows to a primary window, verifies sub-windows with a VLM under four anti-confabulation rules, and fuses evidence with an LLM judge under an evidence-priority hierarchy. Approach B, TMKG: Temporal-Multimodal-Knowledge-Graph, is the contrast: it builds a temporal multimodal knowledge graph, locates a primary cell via graph search, and produces the final answer with a single grounded VLM. SVA reaches a leaderboard accuracy of 0.50 and is our final challenge submission; TMKG reaches 0.35.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuto Kanda, Hayato Tanoue, Takayuki Hori. 2026-05-27. CuriosAI Submission to the CASTLE Challenge at EgoVis 2026. https://arxiv.org/abs/2605.27800
Cite the original work for its findings. Save a collection to share your selection of sources.