arXiv ScienceSearch

arXiv subjects

Jinxing Li

Publications and source records attributed to Jinxing Li.

At least 19 recordsLinked to original sources

Layout-Aware Representation Learning for Open-Set ID Fraud Discovery

Identity-document fraud detection is not a stationary binary classification problem. Adaptive attackers modify templates and fabrication pipelines, making historical fraud labels stale, and successful forgeries recur at scale as coherent campaigns. We therefore study layout-aware representation learning for open-set fraud discovery rather than only closed-set classification. We adapt DINOv3 to the document domain via context-aware SimMIM fine-tuning and supervised metric learning with composite loss that encourages inter-class separability and intra-class compactness. The model is trained with U.S. IDs only. With a lightweight MLP and softmax classifier, the embedding achieves 99.83% layout classification accuracy on Canadian layouts. Moreover, on a dataset of 20,448 Canadian IDs, embedding-space analysis surfaces 276 adaptive physical-fraud cases, including 222 not surfaced by incumbent detectors. The embedding supports similarity-based expansion from a single confirmed seed to additional related cases not linked by conventional metadata graphs. The layout-aware document embeddings provide a production-aligned basis for discovering novel and campaign-scale fraud under distribution shift.

cs.CV

UniFuse: A Unified All-in-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and Misalignments

Current multimodal medical image fusion typically assumes that source images are of high quality and perfectly aligned at the pixel level. Its effectiveness heavily relies on these conditions and often deteriorates when handling misaligned or degraded medical images. To address this, we propose UniFuse, a general fusion framework. By embedding a degradation-aware prompt learning module, UniFuse seamlessly integrates multi-directional information from input images and correlates cross-modal alignment with restoration, enabling joint optimization of both tasks within a unified framework. Additionally, we design an Omni Unified Feature Representation scheme, which leverages Spatial Mamba to encode multi-directional features and mitigate modality differences in feature alignment. To enable simultaneous restoration and fusion within an All-in-One configuration, we propose a Universal Feature Restoration & Fusion module, incorporating the Adaptive LoRA Synergistic Network (ALSN) based on LoRA principles. By leveraging ALSN's adaptive feature representation along with degradation-type guidance, we enable joint restoration and fusion within a single-stage framework. Compared to staged approaches, UniFuse unifies alignment, restoration, and fusion within a single framework. Experimental results across multiple datasets demonstrate the method's effectiveness and significant advantages over existing approaches.

cs.CV

Exploiting hidden singularity on the surface of the Poincar\'e sphere

The classical Pancharatnam-Berry phase, a variant of the geometric phase, arises purely from the modulation of the polarization state of a light beam. Due to its dependence on polarization changes, it cannot be effectively utilized for wavefront shaping in systems that require maintaining a constant (co-polarized) polarization state. Here, we present a novel topologically protected phase modulation mechanism capable of achieving anti-symmetric full 2{\pi} phase shifts with near-unity efficiency for two orthogonal co-polarized channels. Compatible with -- but distinct from- - the dynamic phase, this approach exploits phase circulation around a hidden singularity on the surface of the Poincar\'e sphere. We validate this concept in the microwave regime through the implementation of multi-layer metasurfaces. This new phase modulation mechanism expands the design toolbox of flat optics for light modulation beyond conventional techniques.

physics.optics

Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding

Video Paragraph Grounding (VPG) aims to precisely locate the most appropriate moments within a video that are relevant to a given textual paragraph query. However, existing methods typically rely on large-scale annotated temporal labels and assume that the correspondence between videos and paragraphs is known. This is impractical in real-world applications, as constructing temporal labels requires significant labor costs, and the correspondence is often unknown. To address this issue, we propose a Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding method (DMR-JRG). In this method, retrieval and grounding tasks are mutually reinforced rather than being treated as separate issues. DMR-JRG mainly consists of two branches: a retrieval branch and a grounding branch. The retrieval branch uses inter-video contrastive learning to roughly align the global features of paragraphs and videos, reducing modality differences and constructing a coarse-grained feature space to break free from the need for correspondence between paragraphs and videos. Additionally, this coarse-grained feature space further facilitates the grounding branch in extracting fine-grained contextual representations. In the grounding branch, we achieve precise cross-modal matching and grounding by exploring the consistency between local, global, and temporal dimensions of video segments and textual paragraphs. By synergizing these dimensions, we construct a fine-grained feature space for video and textual features, greatly reducing the need for large-scale annotated temporal labels.

cs.CV

Ring Current Proton Decay Timescales Derived from Van Allen Probe Observations

The Earth's ring current is highly dynamic and is strongly influenced by the solar wind. The ring current alters the planet's magnetic field, defining geomagnetic storms. In this study, we investigate the decay timescales of ring current protons using observations from the Van Allen Probes. Since proton fluxes typically exhibit exponential decay after big storms, the decay time scales are calculated by performing linear regression on the logarithm of the fluxes. We found that in the central region of the ring current, proton decay timescales generally increase with increasing energies and increasing L-shells. The ~10s keV proton decay timescales are about a few days, while the ~100 keV proton decay time scale is about ~10 days, and protons of 269 keV have decay timescales up to ~118 days. These findings provide valuable insights into the ring current dynamics and can contribute to the development of more accurate ring current models.

physics.space-ph

High-g-Factor Phase-Matched Circular Dichroism of Second Harmonic Generation in Chiral Polar Liquids

Circular dichroism is a technologically important phenomenon contrasting the absorption and resultant emission properties between left- and right-handed circularly polarized light. While the chiral handedness of systems mainly determines the mechanism of the circular dichroism in linear optics, the counterpart in the nonlinear optical regime is nontrivial. Here, in contrast to traditional nonlinear circular dichroism responses from structured surfaces, we report on an unprecedented bulk-material-induced circular dichroism of second harmonic generation with a massive g-factor up to 1.8.

physics.optics

Spontaneous electric-polarization topology in confined ferroelectric nematics

Topological spin and polar textures have fascinated people in different areas of physics and technologies. However, the observations are limited in magnetic and solid-state ferroelectric systems. Ferroelectric nematic is the first liquid-state ferroelectric that would carry many possibilities of spatially distributed polarization fields. Contrary to traditional magnetic or crystalline systems, anisotropic liquid crystal interactions can compete with the polarization counterparts, thereby setting a challenge in understating their interplays and the resultant topologies. Here, we discover chiral polarization meron-like structures during the emergence and growth of quasi-2D ferroelectric nematic domains, which are visualized by fluorescence confocal polarizing microscopy and second harmonic generation microscopies. Such micrometre-scale polarization textures are the modified electric variants of the magnetic merons. Unlike the conventional liquid crystal textures driven solely by the elasticity, the polarization field puts additional topological constraints, e.g., head-to-tail asymmetry, to the systems and results in a variety of previously unidentified polar topological patterns. The chirality can emerge spontaneously in polar textures and can be additionally biased by introducing chiral dopants. An extended mean-field modelling for the ferroelectric nematics reveals that the polarization strength of systems plays a dedicated role in determining polarization topology, providing a guide for exploring diverse polar textures in strongly-polarized liquid crystals.

cond-mat.soft

A Survey on Incomplete Multi-view Clustering

Conventional multi-view clustering seeks to partition data into respective groups based on the assumption that all views are fully observed. However, in practical applications, such as disease diagnosis, multimedia analysis, and recommendation system, it is common to observe that not all views of samples are available in many cases, which leads to the failure of the conventional multi-view clustering methods. Clustering on such incomplete multi-view data is referred to as incomplete multi-view clustering. In view of the promising application prospects, the research of incomplete multi-view clustering has noticeable advances in recent years. However, there is no survey to summarize the current progresses and point out the future research directions. To this end, we review the recent studies of incomplete multi-view clustering. Importantly, we provide some frameworks to unify the corresponding incomplete multi-view clustering methods, and make an in-depth comparative analysis for some representative methods from theoretical and experimental perspectives. Finally, some open problems in the incomplete multi-view clustering field are offered for researchers.

cs.LG

Learning Modal-Invariant and Temporal-Memory for Video-based Visible-Infrared Person Re-Identification

Thanks for the cross-modal retrieval techniques, visible-infrared (RGB-IR) person re-identification (Re-ID) is achieved by projecting them into a common space, allowing person Re-ID in 24-hour surveillance systems. However, with respect to the probe-to-gallery, almost all existing RGB-IR based cross-modal person Re-ID methods focus on image-to-image matching, while the video-to-video matching which contains much richer spatial- and temporal-information remains under-explored. In this paper, we primarily study the video-based cross-modal person Re-ID method. To achieve this task, a video-based RGB-IR dataset is constructed, in which 927 valid identities with 463,259 frames and 21,863 tracklets captured by 12 RGB/IR cameras are collected. Based on our constructed dataset, we prove that with the increase of frames in a tracklet, the performance does meet more enhancement, demonstrating the significance of video-to-video matching in RGB-IR person Re-ID. Additionally, a novel method is further proposed, which not only projects two modalities to a modal-invariant subspace, but also extracts the temporal-memory for motion-invariant. Thanks to these two strategies, much better results are achieved on our video-based cross-modal person Re-ID. The code and dataset are released at: https://github.com/VCMproject233/MITML.

cs.CV

HIPA: Hierarchical Patch Transformer for Single Image Super Resolution

Transformer-based architectures start to emerge in single image super resolution (SISR) and have achieved promising performance. Most existing Vision Transformers divide images into the same number of patches with a fixed size, which may not be optimal for restoring patches with different levels of texture richness. This paper presents HIPA, a novel Transformer architecture that progressively recovers the high resolution image using a hierarchical patch partition. Specifically, we build a cascaded model that processes an input image in multiple stages, where we start with tokens with small patch sizes and gradually merge to the full resolution. Such a hierarchical patch mechanism not only explicitly enables feature aggregation at multiple resolutions but also adaptively learns patch-aware features for different image regions, e.g., using a smaller patch for areas with fine details and a larger patch for textureless regions. Meanwhile, a new attention-based position encoding scheme for Transformer is proposed to let the network focus on which tokens should be paid more attention by assigning different weights to different tokens, which is the first time to our best knowledge. Furthermore, we also propose a new multi-reception field attention module to enlarge the convolution reception field from different branches. The experimental results on several public datasets demonstrate the superior performance of the proposed HIPA over previous methods quantitatively and qualitatively.

cs.CV

Pseudocylindrical Convolutions for Learned Omnidirectional Image Compression

Although equirectangular projection (ERP) is a convenient form to store omnidirectional images (also known as 360-degree images), it is neither equal-area nor conformal, thus not friendly to subsequent visual communication. In the context of image compression, ERP will over-sample and deform things and stuff near the poles, making it difficult for perceptually optimal bit allocation. In conventional 360-degree image compression, techniques such as region-wise packing and tiled representation are introduced to alleviate the over-sampling problem, achieving limited success. In this paper, we make one of the first attempts to learn deep neural networks for omnidirectional image compression. We first describe parametric pseudocylindrical representation as a generalization of common pseudocylindrical map projections. A computationally tractable greedy method is presented to determine the (sub)-optimal configuration of the pseudocylindrical representation in terms of a novel proxy objective for rate-distortion performance. We then propose pseudocylindrical convolutions for 360-degree image compression. Under reasonable constraints on the parametric representation, the pseudocylindrical convolution can be efficiently implemented by standard convolution with the so-called pseudocylindrical padding. To demonstrate the feasibility of our idea, we implement an end-to-end 360-degree image compression system, consisting of the learned pseudocylindrical representation, an analysis transform, a non-uniform quantizer, a synthesis transform, and an entropy model. Experimental results on $19,790$ omnidirectional images show that our method achieves consistently better rate-distortion performance than the competing methods. Moreover, the visual quality by our method is significantly improved for all images at all bitrates.

eess.IV

U2-Former: A Nested U-shaped Transformer for Image Restoration

While Transformer has achieved remarkable performance in various high-level vision tasks, it is still challenging to exploit the full potential of Transformer in image restoration. The crux lies in the limited depth of applying Transformer in the typical encoder-decoder framework for image restoration, resulting from heavy self-attention computation load and inefficient communications across different depth (scales) of layers. In this paper, we present a deep and effective Transformer-based network for image restoration, termed as U2-Former, which is able to employ Transformer as the core operation to perform image restoration in a deep encoding and decoding space. Specifically, it leverages the nested U-shaped structure to facilitate the interactions across different layers with different scales of feature maps. Furthermore, we optimize the computational efficiency for the basic Transformer block by introducing a feature-filtering mechanism to compress the token representation. Apart from the typical supervision ways for image restoration, our U2-Former also performs contrastive learning in multiple aspects to further decouple the noise component from the background image. Extensive experiments on various image restoration tasks, including reflection removal, rain streak removal and dehazing respectively, demonstrate the effectiveness of the proposed U2-Former.

cs.CV

BPFNet: A Unified Framework for Bimodal Palmprint Alignment and Fusion

Bimodal palmprint recognition leverages palmprint and palm vein images simultaneously,which achieves high accuracy by multi-model information fusion and has strong anti-falsification property. In the recognition pipeline, the detection of palm and the alignment of region-of-interest (ROI) are two crucial steps for accurate matching. Most existing methods localize palm ROI by keypoint detection algorithms, however the intrinsic difficulties of keypoint detection tasks make the results unsatisfactory. Besides, the ROI alignment and fusion algorithms at image-level are not fully investigaged.To bridge the gap, in this paper, we propose Bimodal Palmprint Fusion Network (BPFNet) which focuses on ROI localization, alignment and bimodal image fusion.BPFNet is an end-to-end framework containing two subnets: The detection network directly regresses the palmprint ROIs based on bounding box prediction and conducts alignment by translation estimation.In the downstream,the bimodal fusion network implements bimodal ROI image fusion leveraging a novel proposed cross-modal selection scheme. To show the effectiveness of BPFNet,we carry out experiments on the large-scale touchless palmprint datasets CUHKSZ-v1 and TongJi and the proposed method achieves state-of-the-art performances.

cs.CV

Dual-Stream Reciprocal Disentanglement Learning for Domain Adaptation Person Re-Identification

Since human-labeled samples are free for the target set, unsupervised person re-identification (Re-ID) has attracted much attention in recent years, by additionally exploiting the source set. However, due to the differences on camera styles, illumination and backgrounds, there exists a large gap between source domain and target domain, introducing a great challenge on cross-domain matching. To tackle this problem, in this paper we propose a novel method named Dual-stream Reciprocal Disentanglement Learning (DRDL), which is quite efficient in learning domain-invariant features. In DRDL, two encoders are first constructed for id-related and id-unrelated feature extractions, which are respectively measured by their associated classifiers. Furthermore, followed by an adversarial learning strategy, both streams reciprocally and positively effect each other, so that the id-related features and id-unrelated features are completely disentangled from a given image, allowing the encoder to be powerful enough to obtain the discriminative but domain-invariant features. In contrast to existing approaches, our proposed method is free from image generation, which not only reduces the computational complexity remarkably, but also removes redundant information from id-related features. Extensive experiments substantiate the superiority of our proposed method compared with the state-of-the-arts. The source code has been released in https://github.com/lhf12278/DRDL.

cs.CV

Observation of Spontaneous Helielectric Nematic Fluids: Electric Analogy to Helimagnets

About a century ago, Born proposed a possible matter of state, ferroelectric fluid, might exist if the dipole moment is strong enough. The experimental realisation of such states needs magnifying molecular polar nature to macroscopic scales in liquids. Here, we report on the discovery of a novel chiral liquid matter state, dubbed chiral ferronematic, stabilized by the local ferroelectric ordering coupled to the chiral helicity. It carries the polar vector rotating helically, corresponding to a helieletric structure, analogous to the magnetic counterpart of helimagnet. The state can be retained down to room-temperature and demonstrates gigantic dielectric and nonlinear optical responses. The novel matter state opens a new chapter for exploring the material space of the diverse ferroelectric liquids.

cond-mat.soft

Touchless Palmprint Recognition based on 3D Gabor Template and Block Feature Refinement

With the growing demand for hand hygiene and convenience of use, palmprint recognition with touchless manner made a great development recently, providing an effective solution for person identification. Despite many efforts that have been devoted to this area, it is still uncertain about the discriminative ability of the contactless palmprint, especially for large-scale datasets. To tackle the problem, in this paper, we build a large-scale touchless palmprint dataset containing 2334 palms from 1167 individuals. To our best knowledge, it is the largest contactless palmprint image benchmark ever collected with regard to the number of individuals and palms. Besides, we propose a novel deep learning framework for touchless palmprint recognition named 3DCPN (3D Convolution Palmprint recognition Network) which leverages 3D convolution to dynamically integrate multiple Gabor features. In 3DCPN, a novel variant of Gabor filter is embedded into the first layer for enhancement of curve feature extraction. With a well-designed ensemble scheme,low-level 3D features are then convolved to extract high-level features. Finally on the top, we set a region-based loss function to strengthen the discriminative ability of both global and local descriptors. To demonstrate the superiority of our method, extensive experiments are conducted on our dataset and other popular databases TongJi and IITD, where the results show the proposed 3DCPN achieves state-of-the-art or comparable performances.

cs.CV

Auto-MVCNN: Neural Architecture Search for Multi-view 3D Shape Recognition

In 3D shape recognition, multi-view based methods leverage human's perspective to analyze 3D shapes and have achieved significant outcomes. Most existing research works in deep learning adopt handcrafted networks as backbones due to their high capacity of feature extraction, and also benefit from ImageNet pretraining. However, whether these network architectures are suitable for 3D analysis or not remains unclear. In this paper, we propose a neural architecture search method named Auto-MVCNN which is particularly designed for optimizing architecture in multi-view 3D shape recognition. Auto-MVCNN extends gradient-based frameworks to process multi-view images, by automatically searching the fusion cell to explore intrinsic correlation among view features. Moreover, we develop an end-to-end scheme to enhance retrieval performance through the trade-off parameter search. Extensive experimental results show that the searched architectures significantly outperform manually designed counterparts in various aspects, and our method achieves state-of-the-art performance at the same time.

cs.CV

Development of polar nematic fluids with giant-\k{appa} dielectric properties

Super-high-\k{appa} materials that exhibit exceptionally high dielectric permittivity are recognized as potential candidates for a wide range of next-generation photonic and electronic devices. Generally, the high dielectricity for achieving a high-\k{appa} state requires a low symmetry of materials so that most of the discovered high-\k{appa} materials are symmetry-broken crystals. There are scarce reports on fluidic high-\k{appa} dielectrics. Here we demonstrate a rational molecular design, supported by machine-learning analyses, that introduces high polarity to asymmetric molecules, successfully realizing super-high-\k{appa} fluid materials (dielectric permittivity, {\epsilon} > 104) and strong second harmonic generation with macroscopic spontaneous polar ordering. The polar structures are confirmed to be identical for all the synthesized materials. Our experiments and computational calculation reveal the unique orientational structures coupled with the emerging polarity. Furthermore, adopting this strategy to high-molecular-weight systems additionally extends the novel material category from monomer to polar polymer materials, creating polar soft matters with spontaneous symmetry breaking.

cond-mat.mtrl-sci