arXiv ScienceSearch

arXiv subjects

Yujia Lin

Publications and source records attributed to Yujia Lin.

5 recordsLinked to original sources

Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization

Ensuring accurate localization of robots in environments without GPS capability is a challenging task. Visual Place Recognition (VPR) techniques can potentially achieve this goal, but existing RGB-based methods are sensitive to changes in illumination, weather, and other seasonal changes. Existing cross-modal localization methods leverage the geometric properties of RGB images and 3D LiDAR maps to reduce the sensitivity issues highlighted above. Currently, state-of-the-art methods struggle in complex scenes, fine-grained or high-resolution matching, and situations where changes can occur in viewpoint. In this work, we introduce a framework we call Semantic-Enhanced Cross-Modal Place Recognition (SCM-PR) that combines high-level semantics utilizing RGB images for robust localization in LiDAR maps. Our proposed method introduces: a VMamba backbone for feature extraction of RGB images; a Semantic-Aware Feature Fusion (SAFF) module for using both place descriptors and segmentation masks; LiDAR descriptors that incorporate both semantics and geometry; and a cross-modal semantic attention mechanism in NetVLAD to improve matching. Incorporating the semantic information also was instrumental in designing a Multi-View Semantic-Geometric Matching and a Semantic Consistency Loss, both in a contrastive learning framework. Our experimental work on the KITTI and KITTI-360 datasets show that SCM-PR achieves state-of-the-art performance compared to other cross-modal place recognition methods.

cs.CV

LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

Text-to-Image (T2I) generation has made significant advancements with diffusion models, yet challenges persist in handling complex instructions, ensuring fine-grained content control, and maintaining deep semantic consistency. Existing T2I models often struggle with tasks like accurate text rendering, precise pose generation, or intricate compositional coherence. Concurrently, Vision-Language Models (LVLMs) have demonstrated powerful capabilities in cross-modal understanding and instruction following. We propose LumiGen, a novel LVLM-enhanced iterative framework designed to elevate T2I model performance, particularly in areas requiring fine-grained control, through a closed-loop, LVLM-driven feedback mechanism. LumiGen comprises an Intelligent Prompt Parsing & Augmentation (IPPA) module for proactive prompt enhancement and an Iterative Visual Feedback & Refinement (IVFR) module, which acts as a "visual critic" to iteratively correct and optimize generated images. Evaluated on the challenging LongBench-T2I Benchmark, LumiGen achieves a superior average score of 3.08, outperforming state-of-the-art baselines. Notably, our framework demonstrates significant improvements in critical dimensions such as text rendering and pose expression, validating the effectiveness of LVLM integration for more controllable and higher-quality image generation.

cs.LG

BCDNet: A Fast Residual Neural Network For Invasive Ductal Carcinoma Detection

It is of great significance to diagnose Invasive Ductal Carcinoma (IDC) in early stage, which is the most common subtype of breast cancer. Although the powerful models in the Computer-Aided Diagnosis (CAD) systems provide promising results, it is still difficult to integrate them into other medical devices or use them without sufficient computation resource. In this paper, we propose BCDNet, which firstly upsamples the input image by the residual block and use smaller convolutional block and a special MLP to learn features. BCDNet is proofed to effectively detect IDC in histopathological RGB images with an average accuracy of 91.6% and reduce training consumption effectively compared to ResNet 50 and ViT-B-16.

eess.IV

Human Digital Twin: A Survey

Digital twin has recently attracted growing attention, leading to intensive research and applications. Along with this, a new research area, dubbed as "human digital twin" (HDT), has emerged. Similar to the conception of digital twin, HDT is referred to as the replica of a physical-world human in the digital world. Nevertheless, HDT is much more complicated and delicate compared to digital twins of any physical systems and processes, due to humans' dynamic and evolutionary nature, including physical, behavioral, social, physiological, psychological, cognitive, and biological dimensions. Studies on HDT are limited, and the research is still in its infancy. In this paper, we first examine the inception, development, and application of the digital twin concept, providing a context within which we formally define and characterize HDT based on the similarities and differences between digital twin and HDT. Then we conduct an extensive literature review on HDT research, analyzing underpinning technologies and establishing typical frameworks in which the core HDT functions or components are organized. Built upon the findings from the above work, we propose a generic architecture for the HDT system and describe the core function blocks and corresponding technologies. Following this, we present the state of the art of HDT technologies and applications in the healthcare, industry, and daily life domain. Finally, we discuss various issues related to the development of HDT and point out the trends and challenges of future HDT research and development.

cs.HC

A Survey on Metaverse: the State-of-the-art, Technologies, Applications, and Challenges

Metaverse is a new type of Internet application and social form that integrates a variety of new technologies. It has the characteristics of multi-technology, sociality, and hyper spatiotemporality. This paper introduces the development status of Metaverse, from the five perspectives of network infrastructure, management technology, basic common technology, virtual reality object connection, and virtual reality convergence, it introduces the technical framework of Metaverse. This paper also introduces the nature of Metaverse's social and hyper spatiotemporality, and discusses the first application areas of Metaverse and some of the problems and challenges it may face.

cs.CY