arXiv · 2512.11234
RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing
Abstract
Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, existing approaches either support a limited input modalities or rely on implicit generation processes that hinder precise control over scene structure and semantics. To address these limitations, we introduce RoomPilot, a unified framework for controllable indoor scene synthesis from multi-modal inputs, including textual descriptions and CAD floor plans. RoomPilot maps heterogeneous inputs into an Indoor Domain-Specific Language (IDSL), which serves as a structured and interpretable semantic representation for describing indoor scenes. Built upon IDSL, RoomPilot presents a hierarchical synthesis pipeline that progressively organizes scenes at the building, room, and object levels, promoting structural coherence and functional consistency across multi-room layouts. Moreover, RoomPilot constructs a curated asset dataset with rich semantic annotations to support high-quality scene synthesis, improving visual realism and appearance consistency. Extensive experiments demonstrate effective multi-modal understanding, fine-grained controllability in scene generation, and improved physical consistency and visual fidelity, marking a significant step toward controllable 3D indoor scene synthesis. Code and model will be available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wentang Chen, Shougao Zhang, Yiman Zhang, Tianhao Zhou, Ruihui Li. 2025-12-12. RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing. https://arxiv.org/abs/2512.11234
Cite the original work for its findings. Save a collection to share your selection of sources.