arXiv ScienceSearch

arXiv subjects

Guohao Zhang

Publications and source records attributed to Guohao Zhang.

12 recordsLinked to original sources

Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM

Monocular panoramic SLAM benefits from substantial visual overlap under large camera rotations, yet remains prone to errors caused by camera tilt, scale drift, and false loop closures. We show that a frozen panoramic geometry foundation model provides useful internal cues beyond its explicit geometric outputs: intermediate tokens encode gravity in the camera frame, while cross-view attention provides a compatibility cue for potential revisits. Building on these cues, we present HALO-SLAM. A gravity readout enables IMU-free spherical upright canonicalization. For loop closure, we introduce a cost-aware three-stage cascade combining DBoW2 event-level retrieval, attention-based compatibility filtering, and dense geometric validation through symmetric submap augmentation. Accepted revisits yield pixel-aligned 3D--3D correspondences in both local gauges, from which robust $\mathrm{Sim}(3)$ constraints are estimated and jointly optimized with sequential constraints in a global pose graph. Across 125 sequences from five real-world panoramic benchmarks, our method achieves \textbf{100\%} sequence success (\textbf{125/125}) under the stated criterion and the lowest ATE among the evaluated methods on all five benchmarks, reducing ATE by \textbf{30--88\%} relative to the best ERP-native baseline on each benchmark.

cs.CV

Large-Aperture All-Solid-State Cascaded Liquid-Crystal Beam Steering for High-Resolution Wide-Field Imaging

High-resolution wide-field imaging is essential for applications requiring simultaneous global coverage and local detail, yet conventional approaches face a fundamental trade-off: wide-FOV cameras sacrifice spatial sampling density by distributing finite detector pixels over a broad angular range, while telephoto systems resolve fine features at the cost of scene coverage. Beam-steering devices can mitigate this trade-off but are currently limited in achieving simultaneously all-solid-state, large aperture, and high-speed operation. Here, we report an all-solid-state large-aperture cascaded liquid-crystal beam-steering (CaLiBS) imaging system that extends the effective angular range of a high-resolution narrow-FOV camera by electrically steering sub-FOVs. The CaLiBS module comprises cascaded liquid crystal waveplates and liquid crystal Pancharatnam-Berry phase gratings; a theoretical voltage-prediction model with a hierarchical search algorithm enables efficient calibration under oblique incidence and 10 times faster calibration speed compared with conventional methods. The calibrated system addresses sub-FOVs across 30.3{\deg} * 30.3{\deg} at 2{\deg} intervals with diffraction efficiency above 60%. Sequential sub-FOV acquisition reconstructs a 34.7 * 34.7 composite image, an 8.6-fold enhancement in spatial-bandwidth product over a single-shot wide-FOV camera using the same detector. Combined with object tracking methods, sub-FOV switching further enables high-resolution tracking of moving vehicles within the wide-area scene. This cascaded LC architecture offers a scalable pathway toward compact, vibration-free, and high-resolution wide-field observation.

physics.optics

Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter

Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision. To address this, we propose an anticipatory risk-guided reinforcement learning framework. Leveraging privileged simulator states, we construct a directionally aligned future collision risk map based on the Closest Point of Approach (CPA). Through an asymmetric actor-critic architecture, the network is trained to self-predict this structured risk, which explicitly guides the visual policy during deployment. A lightweight spatio-temporal encoder extracts motion cues directly from onboard depth sequences, bypassing explicit object tracking or optical flow estimation. Extensive simulated and real-world experiments demonstrate that our method effectively improves safety margins and flight efficiency in dense dynamic clutters compared to existing baselines. Furthermore, the learned policy achieves robust zero-shot Sim-to-Real transfer on a physical quadrotor, relying purely on abstracted spatio-temporal depth sequences and its self-predicted risk priors, validating the effectiveness of our approach and its robust generalization from simulation to reality.

cs.RO

RadioRange: An Open-Source Digital Twin-based Ranging Simulator for UWB, Wi-Fi, and 5G

Accurate RF-based ranging is critical for location-aware wireless systems, yet no open platform exists for fair, reproducible comparison across protocols under realistic hardware impairments. Existing simulators target communication-layer metrics and lack ranging algorithms, impairment models, and positioning-specific evaluation. We present RadioRange, an open-source, positioning-first digital twin that unifies UWB, Wi-Fi, and 5G NR on identical ray-traced physical channels. The platform models eleven independently toggleable hardware impairments across three injection stages, spanning antenna-level offsets, RF circuit non-idealities, and post-compensation CSI residuals, each with documented physical models and protocol-specific defaults. Five first-path ranging detectors and three multipath identification algorithms are provided within a protocol-specific evaluation framework, enabling controlled Monte Carlo benchmarking and systematic ablation studies. The simulator is validated against real-world UWB and Wi-Fi measurements, demonstrating that the channel model captures geometry-dependent multipath bias. RadioRange-Sim is publicly available at https://github.com/Togure/RadioRange.

eess.SP

Functional control of anomalous reflection via engineered metagratings without polarization limitations

Metagratings (MGs) have emerged as a promising platform for manipulating the anomalous propagation of electromagnetic waves. However, traditional methods for designing functional MG-based devices face significant challenges, including complex model structures, time-consuming optimization processes, and specific polarization requirements. In this work, we propose an inverse-design approach to engineer simple MG structures comprising periodic air grooves on a flat metal surface, which can control anomalous reflection without polarization limitations. Through rigorous analytical methods, we derive solutions that achieve perfect retroreflection and perfect specular reflection, thereby leading to functional control over the linearly-polarized electromagnetic waves. Such capabilities enable intriguing functionalities including polarization-dependent retroreflection and polarization-independent retroreflection, as confirmed through full-wave simulations. Our work offers a simple and effective method to control freely electromagnetic waves, with potential applications spanning wavefront engineering, polarization splitting, cloaking technologies, and remote sensing.

physics.optics

Improving GNSS Positioning in Challenging Urban Areas by Digital Twin Database Correction

Accurate positioning technology is the foundation for industry and business applications. Although indoor and outdoor positioning techniques have been well studied separately, positioning performance in the intermediate period of changing the positioning environment is still challenging. This paper proposed a digital twin-aided positioning correction method for seamless positioning focusing on improving the receiver's outdoor positioning performance in urban areas, where the change of the positioning environment usually happens. The proposed algorithm will simulate the positioning solution for virtual receivers in a grid-based digital twin. Based on the simulated positioning solutions, a statistical model will be used to study the positioning characteristics and generate a correction information database for real receivers to improve their positioning performance. This algorithm has a low computation load on the receiver side and does not require a specially designed antenna, making it implementable for small-sized devices.

cs.RO

GNSS Outlier Mitigation Via Graduated Non-Convexity Factor Graph Optimization

Accurate and globally referenced global navigation satellite system (GNSS) based vehicular positioning can be achieved in outlier-free open areas. However, the performance of GNSS can be significantly degraded by outlier measurements, such as multipath effects and non-line-of-sight (NLOS) receptions arising from signal reflections of buildings. Inspired by the advantage of batch historical data in resisting outlier measurements, in this paper, we propose a graduated non-convexity factor graph optimization (FGO-GNC) to improve the GNSS positioning performance, where the impact of GNSS outliers is mitigated by estimating the optimal weightings of GNSS measurements. Different from the existing local solutions, the proposed FGO-GNC employs the non-convex Geman McClure (GM) function to globally estimate the weightings of GNSS measurements via a coarse-to-fine relaxation. The effectiveness of the proposed method is verified through several challenging datasets collected in urban canyons of Hong Kong using automobile level and low-cost smartphone level GNSS receivers.

eess.SP

UrbanLoco: A Full Sensor Suite Dataset for Mapping and Localization in Urban Scenes

Mapping and localization is a critical module of autonomous driving, and significant achievements have been reached in this field. Beyond Global Navigation Satellite System (GNSS), research in point cloud registration, visual feature matching, and inertia navigation has greatly enhanced the accuracy and robustness of mapping and localization in different scenarios. However, highly urbanized scenes are still challenging: LIDAR- and camera-based methods perform poorly with numerous dynamic objects; the GNSS-based solutions experience signal loss and multipath problems; the inertia measurement units (IMU) suffer from drifting. Unfortunately, current public datasets either do not adequately address this urban challenge or do not provide enough sensor information related to mapping and localization. Here we present UrbanLoco: a mapping/localization dataset collected in highly-urbanized environments with a full sensor-suite. The dataset includes 13 trajectories collected in San Francisco and Hong Kong, covering a total length of over 40 kilometers. Our dataset includes a wide variety of urban terrains: urban canyons, bridges, tunnels, sharp turns, etc. More importantly, our dataset includes information from LIDAR, cameras, IMU, and GNSS receivers. Now the dataset is publicly available through the link in the footnote. Dataset Link: https://advdataset2019.wixsite.com/urbanloco.

cs.RO

Measuring the Effects of Scalar and Spherical Colormaps on Ensembles of DMRI Tubes

We report empirical study results on the color encoding of ensemble scalar and orientation to visualize diffusion magnetic resonance imaging (DMRI) tubes. The experiment tested six scalar colormaps for average fractional anisotropy (FA) tasks (grayscale, blackbody, diverging, isoluminant-rainbow, extended-blackbody, and coolwarm) and four three-dimensional (3D) directional encodings for tract tracing tasks (uniform gray, absolute, eigenmap, and Boy's surface embedding). We found that extended-blackbody, coolwarm, and blackbody remain the best three approaches for identifying ensemble average in 3D. Isoluminant-rainbow coloring led to the same ensemble mean accuracy as other colormaps. However, more than 50% of the answers consistently had higher estimates of the ensemble average, independent of the mean values. Hue, not luminance, influences ensemble estimates of mean values. For ensemble orientation-tracing tasks, we found that the Boy's surface embedding (greatest spatial resolution and contrast) and absolute color (lowest spatial resolution and contrast) schemes led to more accurate answers than the eigenmaps scheme (medium resolution and contrast), acting as the uncanny-valley phenomenon of visualization design in terms of accuracy.

cs.GR

Performance Analysis of NDT-based Graph SLAM for Autonomous Vehicle in Diverse Typical Driving Scenarios of Hong Kong

Robust and lane-level positioning is essential for autonomous vehicles. As an irreplaceable sensor, LiDAR can provide continuous and high-frequency pose estimation by means of mapping, on condition that enough environment features are available. The error of mapping can accumulate over time. Therefore, LiDAR is usually integrated with other sensors. In diverse urban scenarios, the environment feature availability relies heavily on the traffic (moving and static objects) and the degree of urbanization. Common LiDAR-based SLAM demonstrations tend to be studied in light traffic and less urbanized area. However, its performance can be severely challenged in deep urbanized cities, such as Hong Kong, Tokyo, and New York with dense traffic and tall buildings. This paper proposes to analyze the performance of standalone NDT-based graph SLAM and its reliability estimation in diverse urban scenarios to further evaluate the relationship between the performance of LiDAR-based SLAM and scenario conditions. The normal distribution transform (NDT) is employed to calculate the transformation between frames of point clouds. Then, the LiDAR odometry is performed based on the calculated continuous transformation. The state-of-the-art graph-based optimization is used to integrate the LiDAR odometry measurements to implement optimization. The 3D building models are generated and the definition of the degree of urbanization based on Skyplot is proposed. Experiments are implemented in different scenarios with different degrees of urbanization and traffic conditions. The results show that the performance of the LiDAR-based SLAM using NDT is strongly related to the traffic condition and degree of urbanization.

cs.RO

Exclusion of GNSS NLOS Receptions Caused by Dynamic Objects in Heavy Traffic Urban Scenarios Using Real-Time 3D Point Cloud: An Approach without 3D Maps

Absolute positioning is an essential factor for the arrival of autonomous driving. Global Navigation Satellites System (GNSS) receiver provides absolute localization for it. GNSS solution can provide satisfactory positioning in open or sub-urban areas, however, its performance suffered in super-urbanized area due to the phenomenon which are well-known as multipath effects and NLOS receptions. The effects dominate GNSS positioning performance in the area. The recent proposed 3D map aided (3DMA) GNSS can mitigate most of the multipath effects and NLOS receptions caused by buildings based on 3D city models. However, the same phenomenon caused by moving objects in urban area is currently not modelled in the 3D geographic information system (GIS). Moving objects with tall height, such as the double-decker bus, can also cause NLOS receptions because of the blockage of GNSS signals by surface of objects. Therefore, we present a novel method to exclude the NLOS receptions caused by double-decker bus in highly urbanized area, Hong Kong. To estimate the geometry dimension and orientation relative to GPS receiver, a Euclidean cluster algorithm and a classification method are used to detect the double-decker buses and calculate their relative locations. To increase the accuracy and reliability of the proposed NLOS exclusion method, an NLOS exclusion criterion is proposed to exclude the blocked satellites considering the elevation, signal noise ratio (SNR) and horizontal dilution of precision (HDOP). Finally, GNSS positioning is estimated by weighted least square (WLS) method using the remaining satellites after the NLOS exclusion. A static experiment was performed near a double-decker bus stop in Hong Kong, which verified the effectiveness of the proposed method.

cs.RO

Overlaying Quantitative Measurement on Networks: An Evaluation of Three Positioning and Nine Visual Marker Techniques

We report results from an experiment on ranking visual markers and node positioning techniques for network visualizations. Inspired by prior ranking studies, we rethink the ranking when the dataset size increases and when the markers are distributed in space. Centrality indices are visualized as node attributes. Our experiment studies nine visual markers and three positioning methods. Our results suggest that direct encoding of quantities improves accuracy by about 20% compared to previous results. Of the three positioning techniques, circular was always in the top group, and matrix and projection switch orders depending on two factors: whether or not the tasks demand symmetry, or the nodes are within closely proximity. Among the most interesting results of ranking the visual markers for comparison tasks are that hue and area fall into the top groups for nearly all multi-scale comparison tasks; Shape (ordered by curvature) is perhaps not as scalable as we have thought and can support more accurate answers only when two quantities are compared; Lightness and slope are least accurate for quantitative comparisons regardless of scale of the comparison tasks. Our experiment is among the first to acquire a complete picture of ranking visual markers in different scales for comparison tasks.

cs.GR