arXiv ScienceSearch

arXiv subjects

Mengjiao Wang

Publications and source records attributed to Mengjiao Wang.

18 recordsLinked to original sources

Generation of dense relativistic electron beams via vortex laser-driven self-generated magnetic pinching

In multi-petawatt laser plasma accelerators, achieving high-density relativistic electron beams is typically accompanied by large transverse divergence, limiting the attainable effective electron density needed for high-flux interaction regimes relevant to laboratory astrophysics. Here we report experimental demonstration of self-generated magnetic pinching (SMP), a collective mechanism that actively regulates transverse beam dynamics using a Laguerre-Gaussian laser at strong relativistic intensity (~8 x 10^19 W/cm^2) interacting with an underdense plasma. The electron beam evolves from a two-lobe high-charge injection structure into a compressed, high-density profile, yielding a threefold reduction in divergence and nearly an order-of-magnitude enhancement in effective beam density compared with a Gaussian driver. Particle-in-cell simulations agree with the experimental observations and reveal that a self-generated azimuthal magnetic field governs the electron dynamics within the SMP regime, which is defined by the forming condition S = 0.717 l a0 [ne(10^18 cm^-3)]^-3/4 = 1, where l, a0, and ne are topological charge, laser amplitude, and plasma density, respectively. A transient kick from a dense inner sheath electron population drives collective magnetic pinching, transforming an initially separated electron distribution into a compressed and well-collimated beam. For higher-power laser systems, the forming condition can be extended to higher plasma densities and larger orbital angular momentum modes, potentially enabling electron beams with charges exceeding several nC and effective densities above 10^19 cm^-3. This mechanism provides a route to overcoming transverse expansion and enhancing rare interaction processes relevant to high-flux particle sources.

physics.plasm-ph

REFA: Real-time Egocentric Facial Animations for Virtual Reality

We present a novel system for real-time tracking of facial expressions using egocentric views captured from a set of infrared cameras embedded in a virtual reality (VR) headset. Our technology facilitates any user to accurately drive the facial expressions of virtual characters in a non-intrusive manner and without the need of a lengthy calibration step. At the core of our system is a distillation based approach to train a machine learning model on heterogeneous data and labels coming form multiple sources, \eg synthetic and real images. As part of our dataset, we collected 18k diverse subjects using a lightweight capture setup consisting of a mobile phone and a custom VR headset with extra cameras. To process this data, we developed a robust differentiable rendering pipeline enabling us to automatically extract facial expression labels. Our system opens up new avenues for communication and expression in virtual environments, with applications in video conferencing, gaming, entertainment, and remote collaboration.

cs.CV

Rigidity for compact hyperbolic complex manifolds

We study the deformation behavior of compact hyperbolic complex manifolds. Let $\pi:\mathcal{X}\rightarrow \Delta$ be a smooth family of compact complex manifolds over the unit disk in $\mathbb{C}$, and $H$ a compact hyperbolic complex manifold. Then the $H$-locus $\{t\in\Delta: X_t\cong H\}$ is either at most a discrete subset of $\Delta$ or the whole $\Delta$. For a smooth family over a compact Riemann surface $Y$, its $H$-locus is either at most finite or the whole $Y$. Furthermore, if $Y$ is isomorphic to $\mathbb{P}^1$ or an elliptic curve, then we conjecture that the $H$-locus is empty or the whole $Y$.

math.CV

A Novel Discrete Memristor-Coupled Heterogeneous Dual-Neuron Model and Its Application in Multi-Scenario Image Encryption

Simulating brain functions using neural networks is an important area of research. Recently, discrete memristor-coupled neurons have attracted significant attention, as memristors effectively mimic synaptic behavior, which is essential for learning and memory. This highlights the biological relevance of such models. This study introduces a discrete memristive heterogeneous dual-neuron network (MHDNN). The stability of the MHDNN is analyzed with respect to initial conditions and a range of neuronal parameters. Numerical simulations demonstrate complex dynamical behaviors. Various neuronal firing patterns are investigated under different coupling strengths, and synchronization phenomena between neurons are explored. The MHDNN is implemented and validated on the STM32 hardware platform. An image encryption algorithm based on the MHDNN is proposed, along with two hardware platforms tailored for multi-scenario police image encryption. These solutions enable real-time and secure transmission of police data in complex environments, reducing hacking risks and enhancing system security.

cs.IR

PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild

This report provides a comprehensive overview of the 4th Pixel-level Video Understanding in the Wild (PVUW) Challenge, held in conjunction with CVPR 2025. It summarizes the challenge outcomes, participating methodologies, and future research directions. The challenge features two tracks: MOSE, which focuses on complex scene video object segmentation, and MeViS, which targets motion-guided, language-based video segmentation. Both tracks introduce new, more challenging datasets designed to better reflect real-world scenarios. Through detailed evaluation and analysis, the challenge offers valuable insights into the current state-of-the-art and emerging trends in complex video segmentation. More information can be found on the workshop website: https://pvuw.github.io/.

cs.CV

FVOS for MOSE Track of 4th PVUW Challenge: 3rd Place Solution

Video Object Segmentation (VOS) is one of the most fundamental and challenging tasks in computer vision and has a wide range of applications. Most existing methods rely on spatiotemporal memory networks to extract frame-level features and have achieved promising results on commonly used datasets. However, these methods often struggle in more complex real-world scenarios. This paper addresses this issue, aiming to achieve accurate segmentation of video objects in challenging scenes. We propose fine-tuning VOS (FVOS), optimizing existing methods for specific datasets through tailored training. Additionally, we introduce a morphological post-processing strategy to address the issue of excessively large gaps between adjacent objects in single-model predictions. Finally, we apply a voting-based fusion method on multi-scale segmentation results to generate the final output. Our approach achieves J&F scores of 76.81% and 83.92% during the validation and testing stages, respectively, securing third place overall in the MOSE Track of the 4th PVUW challenge 2025.

cs.CV

2D Layered Heterojunctions for Photoelectrocatalysis

Two-dimensional (2D) layered nanomaterials heterostructures, arising from the combination of 2D materials with other low-dimensional species, feature large surface area to volume ratio, which provides a high density of active sites for catalytic ap-plications and in particular for (photo)electrocatalysis (PEC). Meanwhile, their unique electronic band structure and high electrical conductivity enable efficient charge transfer (CT) between the active material and the substrate, which is essential for catalytic activity. In recent years, researchers have demonstrated the potential of a range of 2D material interfaces, such as graphene, graphitic carbon nitride (g-C3N4), metal chalcogenides (MCs), and MXenes, for (photo)electrocatalytic applica-tions. For instance, MCs such as MoS2 and WS2 have shown excellent catalytic activity for hydrogen evolution, while gra-phene and MXenes have been used for the reduction of carbon dioxide to higher value chemicals. However, despite their great potential, there are still major challenges that need to be addressed in order to fully realize the potential of 2D materials for PEC. For example, their stability under harsh reaction conditions, as well as their scalability for large-scale production are important factors to be considered. Generating heterojunctions (HJs) by combining 2D layered structures with other na-nomaterials is a promising method to improve the photoelectrocatalytic properties of the former. In this review, we inspect thoroughly the recent literature, to demonstrate the significant potential that arises from utilizing 2D layered heterostructures in PEC processes across a broad spectrum of applications, from energy conversion and storage to environmental remediation. With the ongoing research and development, it is likely that the potential of these materials will be fully expressed in the near future.

cond-mat.mtrl-sci

Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality

Contrastively trained vision-language models have achieved remarkable progress in vision and language representation learning, leading to state-of-the-art models for various downstream multimodal tasks. However, recent research has highlighted severe limitations of these models in their ability to perform compositional reasoning over objects, attributes, and relations. Scene graphs have emerged as an effective way to understand images compositionally. These are graph-structured semantic representations of images that contain objects, their attributes, and relations with other objects in a scene. In this work, we consider the scene graph parsed from text as a proxy for the image scene graph and propose a graph decomposition and augmentation framework along with a coarse-to-fine contrastive learning objective between images and text that aligns sentences of various complexities to the same image. Along with this, we propose novel negative mining techniques in the scene graph space for improving attribute binding and relation understanding. Through extensive experiments, we demonstrate the effectiveness of our approach that significantly improves attribute binding, relation understanding, systematic generalization, and productivity on multiple recently proposed benchmarks (For example, improvements upto $18\%$ for systematic generalization, $16.5\%$ for relation understanding over a strong baseline), while achieving similar or better performance than CLIP on various general multimodal tasks.

cs.CL

Que2Engage: Embedding-based Retrieval for Relevant and Engaging Products at Facebook Marketplace

Embedding-based Retrieval (EBR) in e-commerce search is a powerful search retrieval technique to address semantic matches between search queries and products. However, commercial search engines like Facebook Marketplace Search are complex multi-stage systems optimized for multiple business objectives. At Facebook Marketplace, search retrieval focuses on matching search queries with relevant products, while search ranking puts more emphasis on contextual signals to up-rank the more engaging products. As a result, the end-to-end searcher experience is a function of both relevance and engagement, and the interaction between different stages of the system. This presents challenges to EBR systems in order to optimize for better searcher experiences. In this paper we presents Que2Engage, a search EBR system built towards bridging the gap between retrieval and ranking for end-to-end optimizations. Que2Engage takes a multimodal & multitask approach to infuse contextual information into the retrieval stage and to balance different business objectives. We show the effectiveness of our approach via a multitask evaluation framework and thorough baseline comparisons and ablation studies. Que2Engage is deployed on Facebook Marketplace Search and shows significant improvements in searcher engagement in two weeks of A/B testing.

cs.IR

FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and Captioning

Multimodal tasks in the fashion domain have significant potential for e-commerce, but involve challenging vision-and-language learning problems - e.g., retrieving a fashion item given a reference image plus text feedback from a user. Prior works on multimodal fashion tasks have either been limited by the data in individual benchmarks, or have leveraged generic vision-and-language pre-training but have not taken advantage of the characteristics of fashion data. Additionally, these works have mainly been restricted to multimodal understanding tasks. To address these gaps, we make two key contributions. First, we propose a novel fashion-specific pre-training framework based on weakly-supervised triplets constructed from fashion image-text pairs. We show the triplet-based tasks are an effective addition to standard multimodal pre-training tasks. Second, we propose a flexible decoder-based model architecture capable of both fashion retrieval and captioning tasks. Together, our model design and pre-training approach are competitive on a diverse set of fashion tasks, including cross-modal retrieval, image retrieval with text feedback, image captioning, relative image captioning, and multimodal categorization.

cs.CV

LoopITR: Combining Dual and Cross Encoder Architectures for Image-Text Retrieval

Dual encoders and cross encoders have been widely used for image-text retrieval. Between the two, the dual encoder encodes the image and text independently followed by a dot product, while the cross encoder jointly feeds image and text as the input and performs dense multi-modal fusion. These two architectures are typically modeled separately without interaction. In this work, we propose LoopITR, which combines them in the same network for joint learning. Specifically, we let the dual encoder provide hard negatives to the cross encoder, and use the more discriminative cross encoder to distill its predictions back to the dual encoder. Both steps are efficiently performed together in the same model. Our work centers on empirical analyses of this combined architecture, putting the main focus on the design of the distillation objective. Our experimental results highlight the benefits of training the two encoders in the same network, and demonstrate that distillation can be quite effective with just a few hard negative examples. Experiments on two standard datasets (Flickr30K and COCO) show our approach achieves state-of-the-art dual encoder performance when compared with approaches using a similar amount of data.

cs.CV

Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment

Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require pre-training on a large set of parallel image-text data, which is costly to collect, compared to image-only or text-only data. In this paper, we explore unsupervised Vision-and-Language pre-training (UVLP) to learn the cross-modal representation from non-parallel image and text datasets. We found two key factors that lead to good unsupervised V+L pre-training without parallel data: (i) joint image-and-text input (ii) overall image-text alignment (even for non-parallel data). Accordingly, we propose a novel unsupervised V+L pre-training curriculum for non-parallel texts and images. We first construct a weakly aligned image-text corpus via a retrieval-based approach, then apply a set of multi-granular alignment pre-training tasks, including region-to-tag, region-to-phrase, and image-to-sentence alignment, to bridge the gap between the two modalities. A comprehensive ablation study shows each granularity is helpful to learn a stronger pre-trained model. We adapt our pre-trained model to a set of V+L downstream tasks, including VQA, NLVR2, Visual Entailment, and RefCOCO+. Our model achieves the state-of-art performance in all these tasks under the unsupervised setting.

cs.CV

Control of electronic band profiles through depletion layer engineering in core-shell nanocrystals

The understanding of depletion layers is of major importance to control the optical and electronic properties of metal oxide (MO) nanocrystals (NCs). Here, we show that depletion layer engineering is the main mechanism of photodoping of MO NCs. We show that the introduction of different electronic interfaces induces a double-bending of the electronic bands and a distinct carrier density profile. We found that the light-induced depletion layer modulation and bending of the bands close to the surface of the nanocrystal is the main mechanism responsible for the storage of extra electrons after photodoping in MO NCs. We support our results by a combined experimental and theoretical approach in the case of Sn:In2O3/In2O3 core-shell NCs, in which we compare numerical simulations with empirical modeling and experiments. This allows not only to extract the main mechanism of photodoping in MO NCs but also to engineer the charge storage capability of MO NCs after photodoping. Our results are transferable to other core-multishell systems, opening up a novel direction to control the optoelectronic properties of nanoscale MOs by designing their energetic band profiles through depletion layer engineering.

physics.app-ph

Cryptocurrency Address Clustering and Labeling

Anonymity is one of the most important qualities of blockchain technology. For example, one can simply create a bitcoin address to send and receive funds without providing KYC to any authority. In general, the real identity behind cryptocurrency addresses is not known, however, some addresses can be clustered according to their ownership by analyzing behavioral patterns, allowing those with known attribution to be assigned labels. These labels may be further used for legal and compliance purposes to assist in law enforcement investigations. In this document, we discuss our methodology behind assigning attribution labels to cryptocurrency addresses.

cs.CR

Frequent Item-set Mining without Ubiquitous Items

Frequent Item-set Mining (FIM), sometimes called Market Basket Analysis (MBA) or Association Rule Learning (ARL), are Machine Learning (ML) methods for creating rules from datasets of transactions of items. Most methods identify items likely to appear together in a transaction based on the support (i.e. a minimum number of relative co-occurrence of the items) for that hypothesis. Although this is a good indicator to measure the relevance of the assumption that these items are likely to appear together, the phenomenon of very frequent items, referred to as ubiquitous items, is not addressed in most algorithms. Ubiquitous items have the same entropy as infrequent items, and not contributing significantly to the knowledge. On the other hand, they have strong effect on the performance of the algorithms and sometimes preventing the convergence of the FIM algorithms and thus the provision of meaningful results. This paper discusses the phenomenon of ubiquitous items and demonstrates how ignoring these has a dramatic effect on the computation performances but with a low and controlled effect on the significance of the results.

cs.DS

An Adversarial Neuro-Tensorial Approach For Learning Disentangled Representations

Several factors contribute to the appearance of an object in a visual scene, including pose, illumination, and deformation, among others. Each factor accounts for a source of variability in the data, while the multiplicative interactions of these factors emulate the entangled variability, giving rise to the rich structure of visual object appearance. Disentangling such unobserved factors from visual data is a challenging task, especially when the data have been captured in uncontrolled recording conditions (also referred to as "in-the-wild") and label information is not available. In this paper, we propose the first unsupervised deep learning method (with pseudo-supervision) for disentangling multiple latent factors of variation in face images captured in-the-wild. To this end, we propose a deep latent variable model, where the multiplicative interactions of multiple latent factors of variation are explicitly modelled by means of multilinear (tensor) structure. We demonstrate that the proposed approach indeed learns disentangled representations of facial expressions and pose, which can be used in various applications, including face editing, as well as 3D face reconstruction and classification of facial expression, identity and pose.

cs.CV

A Double Joint Bayesian Approach for J-Vector Based Text-dependent Speaker Verification

J-vector has been proved to be very effective in text-dependent speaker verification with short-duration speech. However, the current state-of-the-art back-end classifiers, e.g. joint Bayesian model, cannot make full use of such deep features. In this paper, we generalize the standard joint Bayesian approach to model the multi-faceted information in the j-vector explicitly and jointly. In our generalization, the j-vector was modeled as a result derived by a generative Double Joint Bayesian (DoJoBa) model, which contains several kinds of latent variables. With DoJoBa, we are able to explicitly build a model that can combine multiple heterogeneous information from the j-vectors. In verification step, we calculated the likelihood to describe whether the two j-vectors having consistent labels or not. On the public RSR2015 data corpus, the experimental results showed that our approach can achieve 0.02\% EER and 0.02\% EER for impostor wrong and impostor correct cases respectively.

cs.SD

Multi-view (Joint) Probability Linear Discrimination Analysis for Multi-view Feature Verification

Multi-view feature has been proved to be very effective in many multimedia applications. However, the current back-end classifiers cannot make full use of such features. In this paper, we propose a method to model the multi-faceted information in the multi-view features explicitly and jointly. In our approach, the feature was modeled as a result derived by a generative multi-view (joint\footnotemark[1]) Probability Linear Discriminant Analysis (PLDA) model, which contains multiple kinds of latent variables. The usual PLDA model only considers one single label. However, in practical use, when using multi-task learned network as feature extractor, the extracted feature are always attached to several labels. This type of feature is called multi-view feature. With multi-view (joint) PLDA, we are able to explicitly build a model that can combine multiple heterogeneous information from the multi-view features. In verification step, we calculated the likelihood to describe whether the two features having consistent labels or not. This likelihood are used in the following decision-making. Experiments have been conducted on large scale verification task. On the public RSR2015 data corpus, the results showed that our approach can achieve 0.02\% EER and 0.09\% EER for impostor wrong and impostor correct cases respectively.

cs.LG