arXiv ScienceSearch

arXiv subjects

Rui Yu

Publications and source records attributed to Rui Yu.

At least 55 records · Page 3Linked to original sources

Enhancing the Travel Experience for People with Visual Impairments through Multimodal Interaction: NaviGPT, A Real-Time AI-Driven Mobile Navigation System

Assistive technologies for people with visual impairments (PVI) have made significant advancements, particularly with the integration of artificial intelligence (AI) and real-time sensor technologies. However, current solutions often require PVI to switch between multiple apps and tools for tasks like image recognition, navigation, and obstacle detection, which can hinder a seamless and efficient user experience. In this paper, we present NaviGPT, a high-fidelity prototype that integrates LiDAR-based obstacle detection, vibration feedback, and large language model (LLM) responses to provide a comprehensive and real-time navigation aid for PVI. Unlike existing applications such as Be My AI and Seeing AI, NaviGPT combines image recognition and contextual navigation guidance into a single system, offering continuous feedback on the user's surroundings without the need for app-switching. Meanwhile, NaviGPT compensates for the response delays of LLM by using location and sensor data, aiming to provide practical and efficient navigation support for PVI in dynamic environments.

cs.HC

Future Does Matter: Boosting 3D Object Detection with Temporal Motion Estimation in Point Cloud Sequences

Accurate and robust LiDAR 3D object detection is essential for comprehensive scene understanding in autonomous driving. Despite its importance, LiDAR detection performance is limited by inherent constraints of point cloud data, particularly under conditions of extended distances and occlusions. Recently, temporal aggregation has been proven to significantly enhance detection accuracy by fusing multi-frame viewpoint information and enriching the spatial representation of objects. In this work, we introduce a novel LiDAR 3D object detection framework, namely LiSTM, to facilitate spatial-temporal feature learning with cross-frame motion forecasting information. We aim to improve the spatial-temporal interpretation capabilities of the LiDAR detector by incorporating a dynamic prior, generated from a non-learnable motion estimation model. Specifically, Motion-Guided Feature Aggregation (MGFA) is proposed to utilize the object trajectory from previous and future motion states to model spatial-temporal correlations into gaussian heatmap over a driving sequence. This motion-based heatmap then guides the temporal feature fusion, enriching the proposed object features. Moreover, we design a Dual Correlation Weighting Module (DCWM) that effectively facilitates the interaction between past and prospective frames through scene- and channel-wise feature abstraction. In the end, a cascade cross-attention-based decoder is employed to refine the 3D prediction. We have conducted experiments on the Waymo and nuScenes datasets to demonstrate that the proposed framework achieves superior 3D detection performance with effective spatial-temporal feature learning.

cs.CV

Emerging Practices for Large Multimodal Model (LMM) Assistance for People with Visual Impairments: Implications for Design

People with visual impairments perceive their environment non-visually and often use AI-powered assistive tools to obtain textual descriptions of visual information. Recent large vision-language model-based AI-powered tools like Be My AI are more capable of understanding users' inquiries in natural language and describing the scene in audible text; however, the extent to which these tools are useful to visually impaired users is currently understudied. This paper aims to fill this gap. Our study with 14 visually impaired users reveals that they are adapting these tools organically -- not only can these tools facilitate complex interactions in household, spatial, and social contexts, but they also act as an extension of users' cognition, as if the cognition were distributed in the visual information. We also found that although the tools are currently not goal-oriented, users accommodate this limitation and embrace the tools' capabilities for broader use. These findings enable us to envision design implications for creating more goal-oriented, real-time processing, and reliable AI-powered assistive technology.

cs.HC

WIA-LD2ND: Wavelet-based Image Alignment for Self-supervised Low-Dose CT Denoising

In clinical examinations and diagnoses, low-dose computed tomography (LDCT) is crucial for minimizing health risks compared with normal-dose computed tomography (NDCT). However, reducing the radiation dose compromises the signal-to-noise ratio, leading to degraded quality of CT images. To address this, we analyze LDCT denoising task based on experimental results from the frequency perspective, and then introduce a novel self-supervised CT image denoising method called WIA-LD2ND, only using NDCT data. The proposed WIA-LD2ND comprises two modules: Wavelet-based Image Alignment (WIA) and Frequency-Aware Multi-scale Loss (FAM). First, WIA is introduced to align NDCT with LDCT by mainly adding noise to the high-frequency components, which is the main difference between LDCT and NDCT. Second, to better capture high-frequency components and detailed information, Frequency-Aware Multi-scale Loss (FAM) is proposed by effectively utilizing multi-scale feature space. Extensive experiments on two public LDCT denoising datasets demonstrate that our WIA-LD2ND, only uses NDCT, outperforms existing several state-of-the-art weakly-supervised and self-supervised methods. Source code is available at https://github.com/zhaohaoyu376/WI-LD2ND.

eess.IV

MoreStyle: Relax Low-frequency Constraint of Fourier-based Image Reconstruction in Generalizable Medical Image Segmentation

The task of single-source domain generalization (SDG) in medical image segmentation is crucial due to frequent domain shifts in clinical image datasets. To address the challenge of poor generalization across different domains, we introduce a Plug-and-Play module for data augmentation called MoreStyle. MoreStyle diversifies image styles by relaxing low-frequency constraints in Fourier space, guiding the image reconstruction network. With the help of adversarial learning, MoreStyle further expands the style range and pinpoints the most intricate style combinations within latent features. To handle significant style variations, we introduce an uncertainty-weighted loss. This loss emphasizes hard-to-classify pixels resulting only from style shifts while mitigating true hard-to-classify pixels in both MoreStyle-generated and original images. Extensive experiments on two widely used benchmarks demonstrate that the proposed MoreStyle effectively helps to achieve good domain generalization ability, and has the potential to further boost the performance of some state-of-the-art SDG methods. Source code is available at https://github.com/zhaohaoyu376/morestyle.

eess.IV

UDQL: Bridging The Gap between MSE Loss and The Optimal Value Function in Offline Reinforcement Learning

The Mean Square Error (MSE) is commonly utilized to estimate the solution of the optimal value function in the vast majority of offline reinforcement learning (RL) models and has achieved outstanding performance. However, we find that its principle can lead to overestimation phenomenon for the value function. In this paper, we first theoretically analyze overestimation phenomenon led by MSE and provide the theoretical upper bound of the overestimated error. Furthermore, to address it, we propose a novel Bellman underestimated operator to counteract overestimation phenomenon and then prove its contraction characteristics. At last, we propose the offline RL algorithm based on underestimated operator and diffusion policy model. Extensive experimental results on D4RL tasks show that our method can outperform state-of-the-art offline RL algorithms, which demonstrates that our theoretical analysis and underestimation way are effective for offline RL tasks.

cs.LG

Is Yang-Mills Theory Unitary in Fractional Spacetime Dimensions?

We present concrete evidence that Yang-Mills theory exhibits non-unitarity in non-integer spacetime dimensions. This violation of unitarity stems from evanescent operators that, while vanishing in four dimensions, are non-zero in general d dimensions. We demonstrate that these evanescent operators lead to the emergence of both negative-norm states and complex anomalous dimensions.

hep-th

Spatial-aware Attention Generative Adversarial Network for Semi-supervised Anomaly Detection in Medical Image

Medical anomaly detection is a critical research area aimed at recognizing abnormal images to aid in diagnosis.Most existing methods adopt synthetic anomalies and image restoration on normal samples to detect anomaly. The unlabeled data consisting of both normal and abnormal data is not well explored. We introduce a novel Spatial-aware Attention Generative Adversarial Network (SAGAN) for one-class semi-supervised generation of health images.Our core insight is the utilization of position encoding and attention to accurately focus on restoring abnormal regions and preserving normal regions. To fully utilize the unlabelled data, SAGAN relaxes the cyclic consistency requirement of the existing unpaired image-to-image conversion methods, and generates high-quality health images corresponding to unlabeled data, guided by the reconstruction of normal images and restoration of pseudo-anomaly images.Subsequently, the discrepancy between the generated healthy image and the original image is utilized as an anomaly score.Extensive experiments on three medical datasets demonstrate that the proposed SAGAN outperforms the state-of-the-art methods.

eess.IV

Light-cone and quasi generalized parton distributions in the 't Hooft model

We present a comprehensive study of the light-cone generalized parton distribution (GPD) and quasi-GPD of a flavor-neutral meson in the 't Hooft model, {\it i.e.}, two-dimensional QCD (\QCDtw) in the $N_c\to\infty$ limit. With the aid of the Hamiltonian approach, we construct the light-cone GPD in terms of the meson's light-cone wave function in the framework of light-front quantization, and express the quasi-GPD in terms of the meson's Bars-Green wave functions and the chiral angle in the framework of equal-time quantization. We show that, both analytically and numerically, the quasi-GPD does approach the light-cone GPD when the meson is boosted to the infinite momentum frame, which justifies the tenet underlying the large momentum effective theory for the off-forward parton distribution. Upon taking the forward limit, the light-cone and quasi-GPDs reduce to the light-cone and quasi-PDFs. As a bonus, we take this chance to correct the incomplete expression of the quasi-PDFs in the 't Hooft model reported in our preceding work [Y. Jia et al. Phys. Rev. D 98, 054011 (2018)].

hep-ph

Progressive Feature Fusion Network for Enhancing Image Quality Assessment

Image compression has been applied in the fields of image storage and video broadcasting. However, it's formidably tough to distinguish the subtle quality differences between those distorted images generated by different algorithms. In this paper, we propose a new image quality assessment framework to decide which image is better in an image group. To capture the subtle differences, a fine-grained network is adopted to acquire multi-scale features. Subsequently, we design a cross subtract block for separating and gathering the information within positive and negative image pairs. Enabling image comparison in feature space. After that, a progressive feature fusion block is designed, which fuses multi-scale features in a novel progressive way. Hierarchical spatial 2D features can thus be processed gradually. Experimental results show that compared with the current mainstream image quality assessment methods, the proposed network can achieve more accurate image quality assessment and ranks second in the benchmark of CLIC in the image perceptual model track.

cs.CV

Magnetic-field-induced electronic instability of Weyl-like fermions in compressed black phosphorus

Revealing the role of Coulomb interaction in topological semimetals with Dirac/Weyl-like band dispersion shapes a new frontier in condensed matter physics. Topological node-line semimetals (TNLSMs), anticipated as a fertile ground for exploring electronic correlation effects due to the anisotropy associated with their node-line structure, have recently attracted considerable attention. In this study, we report an experimental observation for correlation effects in TNLSMs realized by black phosphorus (BP) under hydrostatic pressure. By performing a combination of nuclear magnetic resonance measurements and band calculations on compressed BP, a magnetic-field-induced electronic instability of Weyl-like fermions is identified under an external magnetic field parallel to the so-called nodal ring in the reciprocal space. Anomalous spin fluctuations serving as the fingerprint of electronic instability are observed at low temperatures, and they are observed to maximize at approximately 1.0 GPa. This study presents compressed BP as a realistic material platform for exploring the rich physics in strongly coupled Weyl-like fermions.

cond-mat.str-el

NeRF-Enhanced Outpainting for Faithful Field-of-View Extrapolation

In various applications, such as robotic navigation and remote visual assistance, expanding the field of view (FOV) of the camera proves beneficial for enhancing environmental perception. Unlike image outpainting techniques aimed solely at generating aesthetically pleasing visuals, these applications demand an extended view that faithfully represents the scene. To achieve this, we formulate a new problem of faithful FOV extrapolation that utilizes a set of pre-captured images as prior knowledge of the scene. To address this problem, we present a simple yet effective solution called NeRF-Enhanced Outpainting (NEO) that uses extended-FOV images generated through NeRF to train a scene-specific image outpainting model. To assess the performance of NEO, we conduct comprehensive evaluations on three photorealistic datasets and one real-world dataset. Extensive experiments on the benchmark datasets showcase the robustness and potential of our method in addressing this challenge. We believe our work lays a strong foundation for future exploration within the research community.

cs.CV

Gluonic evanescent operators: two-loop anomalous dimensions

Evanescent operators are a special class of operators that vanish in four-dimensional spacetime but are non-zero in $d=4-2ε$ dimensions. In this paper, we continue our systematic study of the evanescent operators in the pure Yang-Mills theory and focus on their two-loop renormalization. We develop an efficient strategy to compute the two-loop divergences of form factors of high-dimensional and high-length operators by combining the $d$-dimensional unitarity method and the improved tensor reduction techniques. Two-loop anomalous dimensions are obtained for the mass-dimension-10 basis in the planar YM theory, for which both the $\overline{\text{MS}}$ scheme and the finite-renormalization scheme are used. We verify that the two-loop anomalous dimensions are the same in these two schemes at the Wilson-Fisher conformal fixed point. Our computation shows that the evanescent operators are indispensable in order to obtain the correct two-loop anomalous dimensions. This work provides a first computation of the two-loop anomalous dimensions of the complete set of dimension-10 operators. The method we use is also expected to provide an efficient strategy for the two-loop renormalization of general high-dimensional operators.

hep-th

Realization of a Hopf insulator in circuit systems

Three-dimensional (3D) two-band Hopf insulators are a paradigmatic example of topological phases beyond the topological classifications based on powerful methods like $K$-theory and symmetry indicators.Since this class of topological insulating phases was theoretically proposed in 2008, they have attracted significant interest owing to their conceptual novelty, connection to knot theory, and many fascinating physical properties. However, because their realization requires special forms of long-range spin-orbit coupling (SOC), they have not been achieved in any 3D system yet. Here we report the first experimental realization of the long-sought-after Hopf insulator in a 3D circuit system. To implement the Hopf insulator, we construct basic pseudo-spin modules and connection modules that can realize $2\times2$-matrix elements and then design the circuit network according to a tight-binding Hopf insulator Hamiltonian constructed by the Hopf map. By simulating the band structure of the designed circuit network and calculating the Hopf invariant, we find that the circuit realizes a Hopf insulator with Hopf invariant equaling $4$. Experimentally, we measure the band structure of a printed circuit board and find the observed properties of the bulk bands and topological surface states (TSS) are in good agreement with the theoretical predictions, verifying the bulk-boundary correspondence of the Hopf insulator. Our scheme brings the experimental study of Hopf insulators to reality and opens the door to the implementation of more unexplored topological phases beyond the known topological classifications.

cond-mat.mes-hall

Tri-Attention: Explicit Context-Aware Attention Mechanism for Natural Language Processing

In natural language processing (NLP), the context of a word or sentence plays an essential role. Contextual information such as the semantic representation of a passage or historical dialogue forms an essential part of a conversation and a precise understanding of the present phrase or sentence. However, the standard attention mechanisms typically generate weights using query and key but ignore context, forming a Bi-Attention framework, despite their great success in modeling sequence alignment. This Bi-Attention mechanism does not explicitly model the interactions between the contexts, queries and keys of target sequences, missing important contextual information and resulting in poor attention performance. Accordingly, a novel and general triple-attention (Tri-Attention) framework expands the standard Bi-Attention mechanism and explicitly interacts query, key, and context by incorporating context as the third dimension in calculating relevance scores. Four variants of Tri-Attention are generated by expanding the two-dimensional vector-based additive, dot-product, scaled dot-product, and bilinear operations in Bi-Attention to the tensor operations for Tri-Attention. Extensive experiments on three NLP tasks demonstrate that Tri-Attention outperforms about 30 state-of-the-art non-attention, standard Bi-Attention, contextual Bi-Attention approaches and pretrained neural language models1.

cs.CL

Non-abelian gauge fields in circuit systems

Circuits can provide a platform to study novel physics and have been used, for example, to explore various topological phases. Gauge fields-particularly, non-Abelian gauge fields-can play a pivotal role in the design and modulation of novel physical states, but their circuit implementation has so far been limited. Here we show that non-Abelian gauge fields can be synthesized in circuits created from building blocks that consist of capacitors, inductors and resistors. With these building blocks, we create circuit designs for the spin-orbit interaction and the topological Chern state, which are phenomena that represent non-Abelian gauge fields in momentum space. We also use the approach to design non-reciprocal circuits that can be used to implement the non-Abelian Aharonov-Bohm effect in real space.

cond-mat.mes-hall

Gluonic evanescent operators: classification and one-loop renormalization

Evanescent operators are a special class of operators that vanish classically in four-dimensional spacetime, while in general dimensions they are non-zero and are expected to have non-trivial physical effects at the quantum loop level in dimensional regularization. In this paper we initiate the study of evanescent operators in pure Yang-Mills theory. We develop a systematic method for classifying and constructing the $d$-dimensional Lorentz invariant evanescent operators, which start to appear at mass dimension ten. We also compute one-loop form factors for the dimension-ten operators via the $d$-dimensional unitarity method and obtain their one-loop anomalous dimensions. These operators are necessary ingredients in the study of high dimensional operators in effective field theories involving a Yang-Mills sector.

hep-th

Cascade Transformers for End-to-End Person Search

The goal of person search is to localize a target person from a gallery set of scene images, which is extremely challenging due to large scale variations, pose/viewpoint changes, and occlusions. In this paper, we propose the Cascade Occluded Attention Transformer (COAT) for end-to-end person search. Our three-stage cascade design focuses on detecting people in the first stage, while later stages simultaneously and progressively refine the representation for person detection and re-identification. At each stage the occluded attention transformer applies tighter intersection over union thresholds, forcing the network to learn coarse-to-fine pose/scale invariant features. Meanwhile, we calculate each detection's occluded attention to differentiate a person's tokens from other people or the background. In this way, we simulate the effect of other objects occluding a person of interest at the token-level. Through comprehensive experiments, we demonstrate the benefits of our method by achieving state-of-the-art performance on two benchmark datasets.

cs.CV