arXiv ScienceSearch

arXiv subjects

Feiyu Zhao

Publications and source records attributed to Feiyu Zhao.

16 recordsLinked to original sources

AURORA: Active Uncertainty-Driven Re-Orientation for In-Hand Reconstruction

Observing objects grasped by a robot hand is challenging due to severe visual occlusions. Although in-hand manipulation can expose hidden surfaces, existing approaches often rely on predefined or open-loop reorientation strategies that do not explicitly target under-observed regions. We propose AURORA, an active 3D reconstruction framework that closes the loop between online object-centric reconstruction and in-hand reorientation. At its core, Ray-GPIS estimates direction-wise reconstruction uncertainty along candidate viewing rays and selects next-best-view targets using an uncertainty--novelty objective, which are realized through an axis-conditioned in-hand rotation policy. The resulting RGB-D observations are fused incrementally using CAD-free 6D pose tracking and lightweight geometric reconstruction. Experiments demonstrate that AURORA improves reconstruction quality and information-acquisition efficiency over non-active rotation strategies, while Ray-GPIS also outperforms active view-planning baselines in reconstruction performance, action-ranking quality, and planning efficiency. Targeted ablations further validate its robustness to hand occlusion and pose errors. The project webpage is available at https://aurorahand.github.io/

cs.RO

Imaging the 21-cm Signal from the Cosmic Dawn & Epoch of Reionization and the Connection with the Global Signal

The original baseline design for SKA-Low was motivated by the ability to produce tomographic images of the redshifted 21-cm signal, thus allowing the research field to move beyond the simple statistic of the power spectrum. In this chapter we review the imaging capabilities of SKA-Low, the wide variety of methods proposed for quantatively analysing image data, as well as the connection with the global 21-cm signal.

astro-ph.CO

ETac: A Lightweight and Efficient Tactile Simulation Framework for Learning Dexterous Manipulation

Tactile sensors are increasingly integrated into dexterous robotic manipulators to enhance contact perception. However, learning manipulation policies that rely on tactile sensing remains challenging, primarily due to the trade-off between fidelity and computational cost of soft-body simulations. To address this, we present ETac, a tactile simulation framework that models elastomeric soft-body interactions with both high fidelity and efficiency. ETac employs a lightweight data-driven deformation propagation model to capture soft-body contact dynamics, achieving high simulation quality and boosting efficiency that enables large-scale policy training. When serving as the simulation backend, ETac produces surface deformation estimates comparable to FEM and demonstrates applicability for modeling real tactile sensors. Then, we showcase its capability in training a blind grasping policy that leverages large-area tactile feedback to manipulate diverse objects. Running on a single RTX 4090 GPU, ETac supports reinforcement learning across 4,096 parallel environments, achieving a total throughput of 869 FPS. The resulting policy reaches an average success rate of 84.45% across four object types, underscoring ETac's potential to make tactile-based skill learning both efficient and scalable.

cs.RO

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models

Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain. Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth. We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music. HalluAudio comprises over 5K human-verified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA. To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions. Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes. We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music. Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs.

cs.SD

PoiCGAN: A Targeted Poisoning Based on Feature-Label Joint Perturbation in Federated Learning

Federated Learning (FL), as a popular distributed learning paradigm, has shown outstanding performance in improving computational efficiency and protecting data privacy, and is widely applied in industrial image classification. However, due to its distributed nature, FL is vulnerable to threats from malicious clients, with poisoning attacks being a common threat. A major limitation of existing poisoning attack methods is their difficulty in bypassing model performance tests and defense mechanisms based on model anomaly detection. This often results in the detection and removal of poisoned models, which undermines their practical utility. To ensure both the performance of industrial image classification and attacks, we propose a targeted poisoning attack, PoiCGAN, based on feature-label collaborative perturbation. Our method modifies the inputs of the discriminator and generator in the Conditional Generative Adversarial Network (CGAN) to influence the training process, generating an ideal poison generator. This generator not only produces specific poisoned samples but also automatically performs label flipping. Experiments across various datasets show that our method achieves an attack success rate 83.97% higher than baseline methods, with a less than 8.87% reduction in the main task's accuracy. Moreover, the poisoned samples and malicious models exhibit high stealthiness.

cs.LG

NLiPsCalib: An Efficient Calibration Framework for High-Fidelity 3D Reconstruction of Curved Visuotactile Sensors

Recent advances in visuotactile sensors increasingly employ biomimetic curved surfaces to enhance sensorimotor capabilities. Although such curved visuotactile sensors enable more conformal object contact, their perceptual quality is often degraded by non-uniform illumination, which reduces reconstruction accuracy and typically necessitates calibration. Existing calibration methods commonly rely on customized indenters and specialized devices to collect large-scale photometric data, but these processes are expensive and labor-intensive. To overcome these calibration challenges, we present NLiPsCalib, a physics-consistent and efficient calibration framework for curved visuotactile sensors. NLiPsCalib integrates controllable near-field light sources and leverages Near-Light Photometric Stereo (NLiPs) to estimate contact geometry, simplifying calibration to just a few simple contacts with everyday objects. We further introduce NLiPsTac, a controllable-light-source tactile sensor developed to validate our framework. Experimental results demonstrate that our approach enables high-fidelity 3D reconstruction across diverse curved form factors with a simple calibration procedure. We emphasize that our approach lowers the barrier to developing customized visuotactile sensors of diverse geometries, thereby making visuotactile sensing more accessible to the broader community.

cs.RO

An Attempt to Search for Unintended Electromagnetic Radiation from Starlink Satellites with the 21 Centimeter Array: Methodology and RFI Characterization

The rapid expansion of low-Earth-orbit (LEO) megaconstellations introduces new risks to radio astronomy from unintended electromagnetic radiation (UEMR). In this work, we present an attempt to search for UEMR from Starlink satellites using the 21 Centimeter Array (21CMA). Because the sensitivity of a single pod observation is limited, we focus on developing a robust observing and detection pipeline. Using Two-Line Element (TLE) data, we predict satellite transit times to guide the observations, and we define entry into the field of view (FoV) as an apparent declination greater than $85^{\circ}$ with respect to the 21CMA. We analyze the system equivalent flux density (SEFD) and the resulting single-pod sensitivity limits, which explain the detection of emission originating from the ORBCOMM satellites, rather than any detectable broadband UEMR in our dynamic spectra. To validate the methodology, we developed a Python package, orbdemod, to demodulate ORBCOMM downlink signals in our data. The recovered satellite ID agrees with the satellite predicted by our maximum-declination analysis, thereby validating the accuracy of our transit prediction and identification framework. Furthermore, via modulation power spectrum analysis, we show that the impulsive broadband bursts are produced by power line arcing near the array rather than by satellite UEMR.

astro-ph.IM

Simulation and Data Processing of Beamforming Experiments with Four 21CMA Stations

We present an end-to-end simulation and data-processing framework for digital beamforming experiments conducted with four stations of the 21 Centimeter Array (21CMA). Motivated by the need to characterize instrumental systematics, such as those arising from station-level digital beam synthesis and two-stage channelization, and to validate the data-processing pipeline framework for a future upgraded 21CMA with beamforming capability across all stations, we simulate interferometric visibilities using realistic four-station layouts with radio interferometer simulation software OSKAR. Two representative pointings are considered: a bright, complex Cassiopeia A field and a near-north celestial pole (NCP) calibration field. The sky model combines cataloged point sources with a diffuse Galactic component from the Global Sky Model (GSM), and frequency-dependent thermal noise is injected. We further quantify the imprint of two-stage channelization by comparing an ideal beamformer with a coarse-channel phase approximation, demonstrating that off-axis sources exhibit a characteristic piecewise-linear spectral modulation across coarse-channel boundaries. A data-processing pipeline, including Radio Frequency Interference (RFI) mitigation, calibration, imaging, and mosaicking steps consistent with current low-frequency radio astronomy practice, is constructed. The resulting synthetic images and background root-mean-square (RMS) noise measurements demonstrate the feasibility of adapting established 21CMA calibration and imaging strategies to digital beamforming modes, and provide a framework that can be further developed for beam-aware processing in future full-scale 21CMA beamforming observations.

astro-ph.IM

A candidate field for deep imaging of the Epoch of Reionization observed with MWA

Deep imaging of structures from the Cosmic Dawn (CD) and the Epoch of Reionization (EoR) in five targeted fields is one of the highest priority scientific objectives for the Square Kilometre Array (SKA). Selecting 'quiet' fields, which allow deep imaging, is critical for future SKA CD/EoR observations. Pre-observations using existing radio facilities will help estimate the computational capabilities required for optimal data quality and refine data reduction techniques. In this study, we utilize data from the Murchison Widefield Array (MWA) Phase II extended array for a selected field to study the properties of foregrounds. We conduct deep imaging across two frequency bands: 72-103 MHz and 200-231 MHz. We identify up to 2,576 radio sources within a 5-degree radius of the image center (at RA (J2000) $8^h$ , Dec (J2000) 5{\deg}), achieving approximately 80% completeness at 7.7 mJy and 90% at 10.4 mJy for 216 MHz, with a total integration time of 4.43 hours and an average RMS of 1.80 mJy. Additionally, we apply a foreground removal algorithm using Principal Component Analysis (PCA) and calculate the angular power spectra of the residual images. Our results indicate that nearly all resolved radio sources can be successfully removed using PCA, leading to a reduction in foreground power. However, the angular power spectra of the residual map remains over an order of magnitude higher than the theoretically predicted CD/EoR 21 cm signal. Further improvements in data reduction and foreground subtraction techniques will be necessary to enhance these results.

astro-ph.IM

The Short-spacing Interferometer Array for Global 21-cm Signal Detection (SIGMA): Design of the Antennas and Layout

Numerous experiments have been designed to investigate the Cosmic Dawn (CD) and Epoch of Reionization (EoR) by examining redshifted 21-cm emissions from neutral hydrogen. Detecting the global spectrum of redshifted 21-cm signals is typically achieved through single-antenna experiments. However, this global 21-cm signal is deeply embedded in foreground emissions, which are about four orders of magnitude stronger. Extracting this faint signal is a significant challenge, requiring highly precise instrumental calibration. Additionally, accurately modelling receiver noise in single-antenna experiments is inherently complex. An alternative approach using a short-spacing interferometer is expected to alleviate these difficulties because the noise in different receivers is uncorrelated and averages to zero upon cross-correlation. The Short-spacing Interferometer array for Global 21-cm Signal detection (SIGMA) is an upcoming experiment aimed at detecting the global CD/EoR signal using this approach. We describe the SIGMA system with a focus on optimal antenna design and layout, and propose a framework to address cross-talk between antennas in future calibrations. The SIGMA system is intended to serve as a prototype to gain a better understanding of the system's instrumental effects and to optimize its performance further.

astro-ph.IM

Performance of the Segment Anything Model in Various RFI/Events Detection in Radio Astronomy

The emerging era of big data in radio astronomy demands more efficient and higher-quality processing of observational data. While deep learning methods have been applied to tasks such as automatic radio frequency interference (RFI) detection, these methods often face limitations, including dependence on training data and poor generalization, which are also common issues in other deep learning applications within astronomy. In this study, we investigate the use of the open-source image recognition and segmentation model, Segment Anything Model (SAM), and its optimized version, HQ-SAM, due to their impressive generalization capabilities. We evaluate these models across various tasks, including RFI detection and solar radio burst (SRB) identification. For RFI detection, HQ-SAM (SAM) shows performance that is comparable to or even superior to the SumThreshold method, especially with large-area broadband RFI data. In the search for SRBs, HQ-SAM demonstrates strong recognition abilities for Type II and Type III bursts. Overall, with its impressive generalization capability, SAM (HQ-SAM) can be a promising candidate for further optimization and application in RFI and event detection tasks in radio astronomy.

astro-ph.IM

Influence of sources with a spectral peak in the detection of Cosmic Dawn and Epoch of Reionization

Foreground removal is one of the biggest challenges in the detection of the Cosmic Dawn (CD) and Epoch of Reionization (EoR). Various foreground subtraction techniques have been developed based on the spectral smoothness of foregrounds. However, the sources with a spectral peak (SP) at Megahertz may break down the spectral smoothness at low frequencies (< 1000 MHz). In this paper, we cross-match the GaLactic and Extragalactic All-sky Murchison Widefield Array (GLEAM) extragalactic source catalogue with three other radio source catalogues, covering the frequency range from 72 MHz to 1.4 GHz, to search for sources with spectral turnover. 4,423 sources from the GLEAM catalogue are identified as SP sources, representing approximately 3.2 per cent of the GLEAM radio source population. We utilize the properties of SP source candidates obtained from real observations to establish simulations and test the impact of SP sources on the extraction of CD/EoR signals. We statistically compare the differences introduced by SP sources in the residuals after removing the foregrounds with three methods, which are polynomial fitting, Principal Component Analysis (PCA), and fast independent component analysis (FastICA). Our results indicate that the presence of SP sources in the foregrounds has a negligible influence on extracting the CD/EoR signal. After foreground subtraction, the contribution from SP sources to the total power in the two-dimensional (2D) power spectrum within the EoR window is approximately 3 to 4 orders of magnitude lower than the CD/EoR signal.

astro-ph.CO

The Intensity of Diffuse Galactic Emission Reflected by Meteor Trails

We calculate the reflection of diffuse galactic emission by meteor trails and investigate its potential relationship to Meteor Radio Afterglow (MRA). The formula to calculate the reflection of diffuse galactic emission is derived from a simplified case, assuming that the signals are mirrored by the cylindrical over-dense ionization trail of meteors. The overall observed reflection is simulated through a ray tracing algorithm together with the diffuse galactic emission modelled by the GSM sky model. We demonstrate that the spectrum of the reflected signal is broadband and follows a power law with a negative spectral index of around -1.3. The intensity of the reflected signal varies with local sidereal time and the brightness of the meteor and can reach 2000 Jy. These results agree with some previous observations of MRAs. Therefore, we think that the reflection of galactic emission by meteor trails can be a possible mechanism causing MRAs, which is worthy of further research.

astro-ph.IM

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train generalist X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. More details can be found on the project website https://robotics-transformer-x.github.io.

cs.RO

SegRNN: Segment Recurrent Neural Network for Long-Term Time Series Forecasting

RNN-based methods have faced challenges in the Long-term Time Series Forecasting (LTSF) domain when dealing with excessively long look-back windows and forecast horizons. Consequently, the dominance in this domain has shifted towards Transformer, MLP, and CNN approaches. The substantial number of recurrent iterations are the fundamental reasons behind the limitations of RNNs in LTSF. To address these issues, we propose two novel strategies to reduce the number of iterations in RNNs for LTSF tasks: Segment-wise Iterations and Parallel Multi-step Forecasting (PMF). RNNs that combine these strategies, namely SegRNN, significantly reduce the required recurrent iterations for LTSF, resulting in notable improvements in forecast accuracy and inference speed. Extensive experiments demonstrate that SegRNN not only outperforms SOTA Transformer-based models but also reduces runtime and memory usage by more than 78%. These achievements provide strong evidence that RNNs continue to excel in LTSF tasks and encourage further exploration of this domain with more RNN-based approaches. The source code is coming soon.

cs.LG

Detecting HI Galaxies with Deep Neural Networks in the Presence of Radio Frequency Interference

In neutral hydrogen (HI) galaxy survey, a significant challenge is to identify and extract the HI galaxy signal from observational data contaminated by radio frequency interference (RFI). For a drift-scan survey, or more generally a survey of a spatially continuous region, in the time-ordered spectral data, the HI galaxies and RFI all appear as regions which extend an area in the time-frequency waterfall plot, so the extraction of the HI galaxies and RFI from such data can be regarded as an image segmentation problem, and machine learning methods can be applied to solve such problems. In this study, we develop a method to effectively detect and extract signals of HI galaxies based on a Mask R-CNN network combined with the PointRend method. By simulating FAST-observed galaxy signals and potential RFI impacts, we created a realistic data set for the training and testing of our neural network. We compared five different architectures and selected the best-performing one. This architecture successfully performs instance segmentation of HI galaxy signals in the RFI-contaminated time-ordered data (TOD), achieving a precision of 98.64% and a recall of 93.59%.

astro-ph.IM