arXiv ScienceSearch

arXiv subjects

Yang Miao

Publications and source records attributed to Yang Miao.

At least 19 recordsLinked to original sources

Near-Field Velocity Estimation and Doppler-Aware Localization in OFDM Massive MIMO

In Orthogonal Frequency Division Multiplexing (OFDM)-based massive Multiple-Input Multiple-Output (MIMO) near-field (NF) sensing, target motion induces an antenna-dependent bistatic Doppler variation across the array aperture. Ignoring this spatial Doppler variation leads to a model mismatch that degrades NF localization. In this paper, we propose a low-complexity recursive framework for joint radial/transverse velocity estimation and Doppler-aware localization. Initialized by a constant-Doppler coarse localization, the method alternates between closed-form Least Squares Estimator (LSE)-based velocity estimation and antenna-dependent Doppler-aware localization refinement. Simulation and measurement results demonstrate the effectiveness of the proposed framework against two benchmark methods. Compared with a low-complexity constant-Doppler baseline method, the proposed algorithm improves range, angle, and radial velocity estimation results, while also enabling transverse velocity estimation. In the measurement results, the overall localization error decreases from 0.268 m to 0.064 m. The radial and transverse velocity estimation errors are 0.032 m/s and 0.069 m/s, respectively. Compared with a high-complexity exhaustive four-dimensional (4D) Maximum Likelihood Estimator (MLE), the proposed method achieves comparable velocity estimation results while yielding a more accurate localization result when the 4D MLE has a practical finite search grid.

eess.SP

PAWS: Perception of Articulation in the Wild at Scale from Egocentric Videos

Articulation perception aims to recover the motion and structure of articulated objects (e.g., drawers and cupboards), and is fundamental to 3D scene understanding in robotics, simulation, and animation. Existing learning-based methods rely heavily on supervised training with high-quality 3D data and manual annotations, limiting scalability and diversity. To address this limitation, we propose PAWS, a method that directly extracts object articulations from hand-object interactions in large-scale in-the-wild egocentric videos. We evaluate our method on the public data sets, including HD-EPIC and Arti4D data sets, achieving significant improvements over baselines. We further demonstrate that the extracted articulations benefit downstream tasks, including fine-tuning 3D articulation prediction models and enabling robot manipulation. See the project website at https://aaltoml.github.io/PAWS/.

cs.CV

Robust Localization in OFDM-Based Massive MIMO through Phase Offset Calibration

Accurate localization in Orthogonal Frequency Division Multiplexing (OFDM)-based massive Multiple-Input Multiple-Output (MIMO) systems depends critically on phase coherence across subcarriers and antennas. However, practical systems suffer from frequency-dependent and (spatial) antenna-dependent phase offsets, degrading localization accuracy. This paper analytically studies the impact of phase incoherence on localization performance under a static User Equipment (UE) and Line-of-Sight (LoS) scenario. We use two complementary tools. First, we derive the Cram\'er-Rao Lower Bound (CRLB) to quantify the theoretical limits under phase offsets. Then, we develop a Spatial Ambiguity Function (SAF)-based model to characterize ambiguity patterns. Simulation results reveal that spatial phase offsets severely degrade localization performance, while frequency phase offsets have a minor effect in the considered system configuration. To address this, we propose a robust Channel State Information (CSI) calibration framework and validate it using real-world measurements from a practical massive MIMO testbed. The experimental results confirm that the proposed calibration framework significantly improves the localization Root Mean Squared Error (RMSE) from 5 m to 1.2 cm, aligning well with the theoretical predictions.

eess.SP

Experimental Validation of SBFD ISAC in an FR3 Distributed SIMO Testbed

Integrated sensing and communication (ISAC) is a key enabler for future radio networks. This paper presents a sub-band full-duplex (SBFD) ISAC system that assigns non-overlapping OFDM subbands to sensing and communication, enabling simultaneous operation with minimal interference. A distributed testbed with three SIMO nodes is implemented using USRP X410 devices operating at 6.8 GHz with 20 MHz bandwidth per channel. A total of 2048 OFDM subcarriers are partitioned into three subbands: two for sensing using Zadoff-Chu sequences and one for communication using QPSK. Each USRP transmits one subband while receiving signals across all three, forming a 1 x 3 SIMO node. Time synchronization is achieved through host-server coordination without external clock distribution. Indoor measurements, validated against MOCAP ground truth, confirm the feasibility of the SBFD ISAC system. The results demonstrate monostatic sensing with a velocity resolution of 0.145 m/s, and communication under NLoS conditions with a BER of 3.63e-3. Compared with a multiband benchmark requiring three times more spectrum, the SBFD configuration achieves comparable velocity estimation accuracy while conserving resources. The sensing and communication performance trade-off is determined by subcarrier allocation strategy rather than mutual interference.

eess.SP

LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation

We propose LangHOPS, the first Multimodal Large Language Model (MLLM) based framework for open-vocabulary object-part instance segmentation. Given an image, LangHOPS can jointly detect and segment hierarchical object and part instances from open-vocabulary candidate categories. Unlike prior approaches that rely on heuristic or learnable visual grouping, our approach grounds object-part hierarchies in language space. It integrates the MLLM into the object-part parsing pipeline to leverage its rich knowledge and reasoning capabilities, and link multi-granularity concepts within the hierarchies. We evaluate LangHOPS across multiple challenging scenarios, including in-domain and cross-dataset object-part instance segmentation, and zero-shot semantic segmentation. LangHOPS achieves state-of-the-art results, surpassing previous methods by 5.5% Average Precision (AP) (in-domain) and 4.8% (cross-dataset) on the PartImageNet dataset and by 2.5% mIOU on unseen object parts in ADE20K (zero-shot). Ablation studies further validate the effectiveness of the language-grounded hierarchy and MLLM driven part query refinement strategy. The code will be released here.

cs.CV

Dynamic Beamforming and Power Allocation in ISAC via Deep Reinforcement Learning

Integrated Sensing and Communication (ISAC) is a key enabler in 6G networks, where sensing and communication capabilities are designed to complement and enhance each other. One of the main challenges in ISAC lies in resource allocation, which becomes computationally demanding in dynamic environments requiring real-time adaptation. In this paper, we propose a Deep Reinforcement Learning (DRL)-based approach for dynamic beamforming and power allocation in ISAC systems. The DRL agent interacts with the environment and learns optimal strategies through trial and error, guided by predefined rewards. Simulation results show that the DRL-based solution converges within 2000 episodes and achieves up to 80\% of the spectral efficiency of a semidefinite relaxation (SDR) benchmark. More importantly, it offers a significant improvement in runtime performance, achieving decision times of around 20 ms compared to 4500 ms for the SDR method. Furthermore, compared with a Deep Q-Network (DQN) benchmark employing discrete beamforming, the proposed approach achieves approximately 30\% higher sum-rate with comparable runtime. These results highlight the potential of DRL for enabling real-time, high-performance ISAC in dynamic scenarios.

eess.SP

Performance Analysis of Sub-band Full-duplex Cell-free Massive MIMO JCAS Systems

In-band Full-duplex joint communication and sensing systems require self interference cancellation as well as decoupling of the mutual interference between UL communication signals and radar echoes. We present sub-band full-duplex as an alternative duplexing scheme to achieve simultaneous uplink communication and target parameter estimation in a cell-free massive MIMO system. Sub-band full-duplex allows uplink and downlink transmissions simultaneously on non-overlapping frequency resources via explicitly defined uplink and downlink sub-bands in each timeslot. Thus, we propose a sub-band full-duplex cell-free massive MIMO system with active downlink sensing on downlink sub-bands and uplink communication on uplink sub-band. In the proposed system, the target illumination signal is transmitted on the downlink (radar) sub-band whereas uplink users transmit on the uplink (communication) sub-band. By assuming efficient suppression of inter-sub-band interference between radar and communication sub-bands, uplink communication and radar signals can be efficiently processed without mutual interference. We show that each AP can estimate sensing parameters with high accuracy in SBFD cell-free massive MIMO JCAS systems.

eess.SP

Joint Beamforming for Multi-user Multi-target FD ISAC System: A Hybrid GRQ-GA Approach

In this paper, we consider a full-duplex (FD) Integrated Sensing and Communication (ISAC) system, in which the base station (BS) performs downlink and uplink communications with multiple users while simultaneously sensing multiple targets. In the scope of this work, we assume a narrowband and static scenario, aiming to focus on the beamforming and power allocation strategies. We propose a joint beamforming strategy for designing transmit and receive beamformer vectors at the BS. The optimization problem aims to maximize the communication sum-rate, which is critical for ensuring high-quality service to users, while also maintaining accurate sensing performance for detection tasks and adhering to maximum power constraints for efficient resource usage. The optimal receive beamformers are first derived using a closed-form Generalized Rayleigh Quotient (GRQ) solution, reducing the variables to be optimized. Then, the remaining problem is solved using floating-point Genetic Algorithms (GA). The numerical results show that the proposed GA-based solution demonstrates up to a 98% enhancement in sum-rate compared to a baseline half-duplex ISAC system and provides better performance than a benchmark algorithm from the literature. Additionally, it offers insights into sensing performance effects on beam patterns as well as communicationsensing trade-offs in multi-target scenarios.

eess.SP

EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark

Most existing benchmarks for understanding egocentric vision focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first comprehensive benchmark for nighttime egocentric vision, with visual question answering (VQA) as the core task. A key feature of EgoNight is the introduction of day-night aligned videos, which enhance night annotation quality using the daytime data and reveal clear performance gaps between lighting conditions. To achieve this, we collect both synthetic videos rendered by Blender and real-world recordings, ensuring that scenes and actions are visually and temporally aligned. Leveraging these paired videos, we construct EgoNight-VQA, supported by a novel day-augmented night auto-labeling engine and refinement through extensive human verification. Each QA pair is double-checked by annotators for reliability. In total, EgoNight-VQA contains 3658 QA pairs across 90 videos, spanning 12 diverse QA types, with more than 300 hours of human work. Evaluations of state-of-the-art multimodal large language models (MLLMs) reveal substantial performance drops when transferring from day to night, underscoring the challenges of reasoning under low-light conditions. Beyond VQA, EgoNight also introduces two auxiliary tasks, day-night correspondence retrieval and egocentric depth estimation at night, that further explore the boundaries of existing models. We believe EgoNight-VQA provides a strong foundation for advancing application-driven egocentric vision research and for developing models that generalize across illumination domains. The code and data can be found at https://github.com/dehezhang2/EgoNight.

cs.CV

BS-Breath: Respiration Sensing with Cell-free Massive MIMO

This paper demonstrates the feasibility of respiration pattern estimation utilizing a communication-centric cellfree massive MIMO OFDM Base Station (BS). The sensing target is typically positioned near the User Equipment (UE), which transmits uplink pilots to the BS. Our results demonstrate the potential of massive MIMO systems for accurate and reliable vital sign estimation. Initially, we adopt a single antenna sensing solution that combines multiple subcarriers and a breathing projection to align the 2D complex breathing pattern to a single displacement dimension. Then, Weighted Antenna Combining (WAC) aggregates the 1D breathing signals from multiple antennas. The results demonstrate that the combination of space-frequency resources specifically in terms of subcarriers and antennas yields higher accuracy than using only a single antenna or subcarrier. Our results significantly improved respiration estimation accuracy by using multiple subcarriers and antennas. With WAC, we achieved an average correlation of 0.8 with ground truth data, compared to 0.6 for single antenna or subcarrier methods, a 0.2 correlation increase. Moreover, the system produced perfect breathing rate estimates. These findings suggest that the limited bandwidth (18 MHz in the testbed) can be effectively compensated by utilizing spatial resources, such as distributed antennas.

eess.SP

COST INTERACT Whitepaper on Signal Processing for Communications, Localization, and Intergrated Sensing and Communication

The upcoming next generation of wireless communication is anticipated to revolutionize the conventional functionalities of the network by adding sensing and localization capabilities, low-power communication, wireless brain computer interactions, massive robotics and autonomous systems connection. Furthermore, the key performance indicators expected for the 6G of mobile communications promise challenging operating conditions, such as user data rates of 1 Tbps, end-to-end latency of less than 1 ms, and vehicle speeds of 1000 km per hour. This evolution needs new techniques, not only to improve communications, but also to provide localization and sensing with an efficient use of the radio resources. The goal of INTERACT Working Group 2 is to design novel physical layer technologies that can meet these KPI, by combining the data information from statistical learning with the theoretical knowledge of the transmitted signal structure. Waveforms and coding, advanced multiple-input multiple-output and all the required signal processing, in sub-6-GHz, millimeter-wave bands and upper-mid-band, are considered while aiming at designing these new communications, positioning and localization techniques. This White Paper summarizes our main approaches and contributions.

eess.SP

User-Movement-Robust Virtual Reality Through Dual-Beam Reception in mmWave Networks

Utilizing the mmWave band can potentially achieve the high data rate needed for realistic and seamless interaction within a virtual reality (VR) application. To this end, beamforming in both the access point (AP) and head-mounted display (HMD) sides is necessary. The main challenge in this use case is the specific and highly dynamic user movement, which causes beam misalignment, degrading the received signal level and potentially leading to outages. This study examines mmWave-based coordinated multi-point networks for VR applications, where two or multiple APs cooperatively transmit the signals to an HMD for connectivity diversity. Instead of using omnireception, we propose dual-beam reception based on the analog beamforming at the HMD, enhancing the receive beamforming gain towards serving APs while achieving diversity. Evaluation using actual HMD movement data demonstrates the effectiveness of our approach, showcasing a reduction in outage rates of up to 13% compared to quasi-omnidirectional reception with two serving APs, and a 17% decrease compared to steerable single-beam reception with a serving AP. Widening the separation angle between two APs can further reduce outage rates due to head rotation as rotations can still be tracked using the steerable multi-beam, albeit at the expense of received signal levels reduction during the non-outage period.

eess.SP

Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description

3D scene understanding is a long-standing challenge in computer vision and a key component in enabling mixed reality, wearable computing, and embodied AI. Providing a solution to these applications requires a multifaceted approach that covers scene-centric, object-centric, as well as interaction-centric capabilities. While there exist numerous datasets and algorithms approaching the former two problems, the task of understanding interactable and articulated objects is underrepresented and only partly covered in the research field. In this work, we address this shortcoming by introducing: (1) Articulate3D, an expertly curated 3D dataset featuring high-quality manual annotations on 280 indoor scenes. Articulate3D provides 8 types of annotations for articulated objects, covering parts and detailed motion information, all stored in a standardized scene representation format designed for scalable 3D content creation, exchange and seamless integration into simulation environments. (2) USDNet, a novel unified framework capable of simultaneously predicting part segmentation along with a full specification of motion attributes for articulated objects. We evaluate USDNet on Articulate3D as well as two existing datasets, demonstrating the advantage of our unified dense prediction approach. Furthermore, we highlight the value of Articulate3D through cross-dataset and cross-domain evaluations and showcase its applicability in downstream tasks such as scene editing through LLM prompting and robotic policy training for articulated object manipulation. We provide open access to our dataset, benchmark, and method's source code.

cs.CV

Measurement-based Characterization of ISAC Channels with Distributed Beamforming at Dual mmWave Bands and with Human Body Scattering and Blockage

In this paper, we introduce our millimeter-wave (mmWave) radio channel measurement for integrated sensing and communication (ISAC) scenarios with distributed links at dual bands in an indoor cavity; we also characterize the channel in delay and azimuth-angular domains for the scenarios with the presence of 1 person with varying locations and facing orientations. In our setting of distributed links with two transmitters and two receivers where each transceiver operates at two bands, we can measure two links whose each transmitter faces to one receiver and thus capable of line-of-sight (LOS) communication; these two links have crossing Fresnel zones. We have another two links capable of capturing the reflectivity from the target presenting in the test area (as well as the background). The numerical results in this paper focus on analyzing the channel with the presence of one person. It is evident that not only the human location, but also the human facing orientation, shall be taken into account when modeling the ISAC channel.

eess.SP

COST CA20120 INTERACT Framework of Artificial Intelligence Based Channel Modeling

Accurate channel models are the prerequisite for communication-theoretic investigations as well as system design. Channel modeling generally relies on statistical and deterministic approaches. However, there are still significant limits for the traditional modeling methods in terms of accuracy, generalization ability, and computational complexity. The fundamental reason is that establishing a quantified and accurate mapping between physical environment and channel characteristics becomes increasing challenging for modern communication systems. Here, in the context of COST CA20120 Action, we evaluate and discuss the feasibility and implementation of using artificial intelligence (AI) for channel modeling, and explore where the future of this field lies. Firstly, we present a framework of AI-based channel modeling to characterize complex wireless channels. Then, we highlight in detail some major challenges and present the possible solutions: i) estimating the uncertainty of AI-based channel predictions, ii) integrating prior knowledge of propagation to improve generalization capabilities, and iii) interpretable AI for channel modeling. We present and discuss illustrative numerical results to showcase the capabilities of AI-based channel modeling.

cs.IT

SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs

We introduce a novel problem, i.e., the localization of an input image within a multi-modal reference map represented by a database of 3D scene graphs. These graphs comprise multiple modalities, including object-level point clouds, images, attributes, and relationships between objects, offering a lightweight and efficient alternative to conventional methods that rely on extensive image databases. Given the available modalities, the proposed method SceneGraphLoc learns a fixed-sized embedding for each node (i.e., representing an object instance) in the scene graph, enabling effective matching with the objects visible in the input query image. This strategy significantly outperforms other cross-modal methods, even without incorporating images into the map embeddings. When images are leveraged, SceneGraphLoc achieves performance close to that of state-of-the-art techniques depending on large image databases, while requiring three orders-of-magnitude less storage and operating orders-of-magnitude faster. The code will be made public.

cs.CV

Vital Signs Estimation Using a 26 GHz Multi-Beam Communication Testbed

This paper presents a novel pipeline for vital sign monitoring using a 26 GHz multi-beam communication testbed. In context of Joint Communication and Sensing (JCAS), the advanced communication capability at millimeter-wave bands is comparable to the radio resource of radars and is promising to sense the surrounding environment. Being able to communicate and sense the vital sign of humans present in the environment will enable new vertical services of telecommunication, i.e., remote health monitoring. The proposed processing pipeline leverages spatially orthogonal beams to estimate the vital sign - breath rate and heart rate - of single and multiple persons in static scenarios from the raw Channel State Information samples. We consider both monostatic and bistatic sensing scenarios. For monostatic scenario, we employ the phase time-frequency calibration and Discrete Wavelet Transform to improve the performance compared to the conventional Fast Fourier Transform based methods. For bistatic scenario, we use K-means clustering algorithm to extract multi-person vital signs due to the distinct frequency-domain signal feature between single and multi-person scenarios. The results show that the estimated breath rate and heart rate reach below 2 beats per minute (bpm) error compared to the reference captured by on-body sensor for the single-person monostatic sensing scenario with body-transceiver distance up to 2 m, and the two-person bistatic sensing scenario with BS-UE distance up to 4 m. The presented work does not optimize the OFDM waveform parameters for sensing; it demonstrates a promising JCAS proof-of-concept in contact-free vital sign monitoring using mmWave multi-beam communication systems.

eess.SP

Volumetric Semantically Consistent 3D Panoptic Mapping

We introduce an online 2D-to-3D semantic instance mapping algorithm aimed at generating comprehensive, accurate, and efficient semantic 3D maps suitable for autonomous agents in unstructured environments. The proposed approach is based on a Voxel-TSDF representation used in recent algorithms. It introduces novel ways of integrating semantic prediction confidence during mapping, producing semantic and instance-consistent 3D regions. Further improvements are achieved by graph optimization-based semantic labeling and instance refinement. The proposed method achieves accuracy superior to the state of the art on public large-scale datasets, improving on a number of widely used metrics. We also highlight a downfall in the evaluation of recent studies: using the ground truth trajectory as input instead of a SLAM-estimated one substantially affects the accuracy, creating a large gap between the reported results and the actual performance on real-world data.

cs.RO