arXiv · 2609.26360
Hierarchical Floorplan-Guided Vision-Language Exploration for Embodied Question Answering
Abstract
Embodied Question Answering (EQA) requires an agent to explore a previously unseen environment, gather relevant information, and answer questions about the scene. Recent approaches leverage Vision-Language Models (VLMs) together with semantic maps or scene graphs to guide exploration. However, exploration is typically driven only by local observations, while structural priors about the environment remain largely unused. We propose HFLEX-EQA, a hierarchical EQA framework that combines online scene graph construction, VLM- based planning, semantic frontier exploration, and floorplan priors. The system incrementally builds a hierarchical scene graph and an open-vocabulary occupancy map from RGB-D observations, enabling a VLM to jointly reason over the scene graph, task-relevant visual observations, exploration history, and an estimated topological floorplan. Furthermore, we introduce a room-discovery strategy that leverages the floorplan and open-vocabulary frontier semantics to guide exploration toward semantically relevant yet currently unobserved room types. We evaluate HFLEX-EQA on the OpenEQA and ExploreEQA benchmarks and demonstrate deployment on a quadruped robot in real indoor environments. Our results demonstrate the benefit of combining VLM-based hierarchical planning with structural floorplan priors for the EQA task.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Albert Gassol Puigjaner, Kostas Alexis. 2026-09-22. Hierarchical Floorplan-Guided Vision-Language Exploration for Embodied Question Answering. https://arxiv.org/abs/2609.26360
Cite the original work for its findings. Save a collection to share your selection of sources.