DoRF++: Spherical Representation Learning over Doppler Radiance Fields for Robust Wi-Fi Sensing
Motivated by the IEEE 802.11bf effort to standardize advanced WLAN sensing, interest in Wi-Fi Channel State Information (CSI) for passive, device-free, and privacy-preserving activity and gesture recognition has grown rapidly. Recent studies have shown that Doppler velocity projections extracted from CSI, which directly reflect human-motion velocity, enable more robust human activity recognition (HAR) and stronger generalization across users and unseen conditions. Nevertheless, reliable generalization under real-world variability remains a major challenge that hinders the adoption of Wi-Fi sensing. To address this challenge, we build on Doppler Radiance Fields (DoRF), adapting the concept of neural radiance fields (NeRF) from computer vision to Wi-Fi sensing. DoRF treats the extracted Doppler velocity projections as sparse virtual-camera views of human motion. From these views it recovers a regularized latent 3-D motion descriptor whose projections along learned effective Doppler directions explain the CSI observations. The recovered motion is then re-projected onto an equiangular grid of directions on the unit sphere, yielding a spherical representation of the underlying activity. Because DoRF produces a signal defined on the sphere, this paper introduces DoRF++, a spherical representation-learning model that applies quadrature-aware spherical Transformers directly to the spherical field for activity classification. Experiments on our collected hand-gesture dataset show that DoRF++ noticeably outperforms state-of-the-art Wi-Fi-based HAR methods in cross-user generalization accuracy, especially for difficult gestures when only a single multi-antenna receiver access point (AP) is available.