arXiv · 2609.23372
Vision-Wireless Fusion for Multi-User Localization: A Cross-Modal Transformer Approach
Abstract
Accurate multi-user localization is challenging in complex urban environments, where wireless measurements can become ambiguous under noise, blockage, and multipath, while visual observations provide complementary spatial context. This paper presents a vision-wireless fusion framework for multi-user localization using pilot-indexed channel state information (CSI). Orthogonal pilot indices preserve the identities of the communicating UEs in the CSI token sequence and localization outputs. The model encodes each pilot-indexed CSI observation as a query token and uses cross-attention to retrieve user-specific information from spatial visual memory. Self-attention among CSI tokens further captures inter-user interactions, while the resulting multimodal representations are used for user-wise localization. Experiments on different datasets show consistent improvements over model-based, CSI-only, and multimodal-fusion baselines. Further experiments evaluate the model under different wireless and visual conditions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Can Zheng, Jiguang He, Guofa Cai, Henk Wymeersch, Merouane Debbah. 2026-09-20. Vision-Wireless Fusion for Multi-User Localization: A Cross-Modal Transformer Approach. https://arxiv.org/abs/2609.23372
Cite the original work for its findings. Save a collection to share your selection of sources.