arXiv Science⌕ Search

arXiv subjects

Ping Liu

Publications and source records attributed to Ping Liu.

At least 91 records · Page 5Linked to original sources

LiGNN: Graph Neural Networks at LinkedIn

In this paper, we present LiGNN, a deployed large-scale Graph Neural Networks (GNNs) Framework. We share our insight on developing and deployment of GNNs at large scale at LinkedIn. We present a set of algorithmic improvements to the quality of GNN representation learning including temporal graph architectures with long term losses, effective cold start solutions via graph densification, ID embeddings and multi-hop neighbor sampling. We explain how we built and sped up by 7x our large-scale training on LinkedIn graphs with adaptive sampling of neighbors, grouping and slicing of training data batches, specialized shared-memory queue and local gradient optimization. We summarize our deployment lessons and learnings gathered from A/B test experiments. The techniques presented in this work have contributed to an approximate relative improvements of 1% of Job application hearing back rate, 2% Ads CTR lift, 0.5% of Feed engaged daily active users, 0.2% session lift and 0.1% weekly active user lift from people recommendation. We believe that this work can provide practical solutions and insights for engineers who are interested in applying Graph neural networks at large scale.

cs.LG↗

Spectra and pseudo-spectra of tridiagonal $k$-Toeplitz matrices and the topological origin of the non-Hermitian skin effect

We establish new results on the spectra and pseudo-spectra of tridiagonal $k$-Toeplitz operators and matrices. In particular, we prove the connection between the winding number of the eigenvalues of the symbol function and the exponential decay of the associated eigenvectors (or pseudo-eigenvectors). Our results elucidate the topological origin of the non-Hermitian skin effect in general one-dimensional polymer systems of subwavelength resonators with imaginary gauge potentials, proving the observation and conjecture in arXiv:2307.13551. We also numerically verify our theory for these systems.

math-ph↗

Unified Diffusion-Based Rigid and Non-Rigid Editing with Text and Image Guidance

Existing text-to-image editing methods tend to excel either in rigid or non-rigid editing but encounter challenges when combining both, resulting in misaligned outputs with the provided text prompts. In addition, integrating reference images for control remains challenging. To address these issues, we present a versatile image editing framework capable of executing both rigid and non-rigid edits, guided by either textual prompts or reference images. We leverage a dual-path injection scheme to handle diverse editing scenarios and introduce an integrated self-attention mechanism for fusion of appearance and structural information. To mitigate potential visual artifacts, we further employ latent fusion techniques to adjust intermediate latents. Compared to previous work, our approach represents a significant advance in achieving precise and versatile image editing. Comprehensive experiments validate the efficacy of our method, showcasing competitive or superior results in text-based editing and appearance transfer tasks, encompassing both rigid and non-rigid settings.

cs.CV↗

Rethinking the Up-Sampling Operations in CNN-based Generative Network for Generalizable Deepfake Detection

Recently, the proliferation of highly realistic synthetic images, facilitated through a variety of GANs and Diffusions, has significantly heightened the susceptibility to misuse. While the primary focus of deepfake detection has traditionally centered on the design of detection algorithms, an investigative inquiry into the generator architectures has remained conspicuously absent in recent years. This paper contributes to this lacuna by rethinking the architectures of CNN-based generators, thereby establishing a generalized representation of synthetic artifacts. Our findings illuminate that the up-sampling operator can, beyond frequency-based artifacts, produce generalized forgery artifacts. In particular, the local interdependence among image pixels caused by upsampling operators is significantly demonstrated in synthetic images generated by GAN or diffusion. Building upon this observation, we introduce the concept of Neighboring Pixel Relationships(NPR) as a means to capture and characterize the generalized structural artifacts stemming from up-sampling operations. A comprehensive analysis is conducted on an open-world dataset, comprising samples generated by \tft{28 distinct generative models}. This analysis culminates in the establishment of a novel state-of-the-art performance, showcasing a remarkable \tft{11.6\%} improvement over existing methods. The code is available at https://github.com/chuangchuangtan/NPR-DeepfakeDetection.

cs.CV↗

The Non-Hermitian Skin Effect With Three-Dimensional Long-Range Coupling

We study the non-Hermitian skin effect in a three-dimensional system of finitely many subwavelength resonators with an imaginary gauge potential. We introduce a discrete approximation of the eigenmodes and eigenfrequencies of the system in terms of the eigenvectors and eigenvalues of the so-called gauge capacitance matrix $\mathcal{C}_N^γ$, which is a dense matrix due to long-range interactions in the system. Based on translational invariance of this matrix and the decay of its off-diagonal entries, we prove the condensation of the eigenmodes at one edge of the structure by showing the exponential decay of its pseudo-eigenvectors. In particular, we consider a range-k approximation to keep the long-range interaction to a certain extent, thus obtaining a k-banded gauge capacitance matrix $\mathcal{C}_{N,k}^γ$ . Using techniques for Toeplitz matrices and operators, we establish the exponential decay of the pseudo-eigenvectors of $\mathcal{C}_{N,k}^γ$ and demonstrate that they approximate those of the gauge capacitance matrix $\mathcal{C}_N^γ$ well. Our results are numerically verified. In particular, we show that long-range interactions affect only the first eigenmodes in the system. As a result, a tridiagonal approximation of the gauge capacitance matrix, similar to the nearest-neighbour approximation in quantum mechanics, provides a good approximation for the higher modes. Moreover, we also illustrate numerically the behaviour of the eigenmodes and the stability of the non-Hermitian skin effect with respect to disorder in a variety of three-dimensional structures.

math.AP↗

CLE Diffusion: Controllable Light Enhancement Diffusion Model

Low light enhancement has gained increasing importance with the rapid development of visual creation and editing. However, most existing enhancement algorithms are designed to homogeneously increase the brightness of images to a pre-defined extent, limiting the user experience. To address this issue, we propose Controllable Light Enhancement Diffusion Model, dubbed CLE Diffusion, a novel diffusion framework to provide users with rich controllability. Built with a conditional diffusion model, we introduce an illumination embedding to let users control their desired brightness level. Additionally, we incorporate the Segment-Anything Model (SAM) to enable user-friendly region controllability, where users can click on objects to specify the regions they wish to enhance. Extensive experiments demonstrate that CLE Diffusion achieves competitive performance regarding quantitative metrics, qualitative results, and versatile controllability. Project page: https://yuyangyin.github.io/CLEDiffusion/

cs.CV↗

Stability of the non-Hermitian skin effect

This paper shows that the skin effect in systems of non-Hermitian subwavelength resonators is robust with respect to random imperfections in the system. The subwavelength resonators are highly contrasting material inclusions that resonate in a low-frequency regime. The non-Hermiticity is due to the introduction of an imaginary gauge potential, which leads to a skin effect that is manifested by the system's eigenmodes accumulating at one edge of the structure. We elucidate the topological protection of the associated (real) eigenfrequencies and illustrate the competition between the two different localisation effects present when the system is randomly perturbed: the non-Hermitian skin effect and the disorder-induced Anderson localisation. We show that, as the strength of the disorder increases, more and more eigenmodes become localised in the bulk. Our results are based on an asymptotic matrix model for subwavelength physics and can be generalised also to tight-binding models in condensed matter theory.

math-ph↗

Dual Stage Stylization Modulation for Domain Generalized Semantic Segmentation

Obtaining sufficient labeled data for training deep models is often challenging in real-life applications. To address this issue, we propose a novel solution for single-source domain generalized semantic segmentation. Recent approaches have explored data diversity enhancement using hallucination techniques. However, excessive hallucination can degrade performance, particularly for imbalanced datasets. As shown in our experiments, minority classes are more susceptible to performance reduction due to hallucination compared to majority classes. To tackle this challenge, we introduce a dual-stage Feature Transform (dFT) layer within the Adversarial Semantic Hallucination+ (ASH+) framework. The ASH+ framework performs a dual-stage manipulation of hallucination strength. By leveraging semantic information for each pixel, our approach adaptively adjusts the pixel-wise hallucination strength, thus providing fine-grained control over hallucination. We validate the effectiveness of our proposed method through comprehensive experiments on publicly available semantic segmentation benchmark datasets (Cityscapes and SYNTHIA). Quantitative and qualitative comparisons demonstrate that our approach is competitive with state-of-the-art methods for the Cityscapes dataset and surpasses existing solutions for the SYNTHIA dataset. Code for our framework will be made readily available to the research community.

cs.CV↗

A mathematical theory of resolution limits for super-resolution of positive sources

A priori information on the positivity of source intensities is ubiquitous in imaging fields and is also important for a multitude of super-resolution and deconvolution algorithms. However, the fundamental resolution limit of positive sources is still unknown, and research in this field is very limited indeed. In this work, we analyze the super-resolving capacity for number and location recoveries in the super-resolution of positive sources and aim to answer the resolution limit problem in a rigorous manner. Specifically, we introduce the computational resolution limit for respectively the number detection and location recovery in the one-dimensional super-resolution problem and quantitatively characterize their dependency on the cutoff frequency, signal-to-noise ratio, and the sparsity of the sources. As a direct consequence, we show that targeting at the sparest positive solution in the super-resolution already provides the optimal resolution order. These results are generalized to multi-dimensional spaces. Our estimates indicate that there exist phase transitions in the corresponding reconstructions, which are confirmed by numerical experiments. On the other hand, despite the fact that positivity plays important roles in improving the resolution of certain super-resolution algorithms, our theory has made several different but significant discoveries: i) The a priori information of positivity cannot further improve the order of the resolution limit; ii) The positivity of the source sometimes deteriorates the resolution limit instead of enhancing it. In particular, under certain signal-to-noise ratio, two point sources with different phases actually have a better resolution limit than those with the same one.

eess.IV↗

Perturbed Block Toeplitz matrices and the non-Hermitian skin effect in dimer systems of subwavelength resonators

The aim of this paper is fourfold: (i) to obtain explicit formulas for the eigenpairs of perturbed tridiagonal block Toeplitz matrices; (ii) to make use of such formulas in order to provide a mathematical justification of the non-Hermitian skin effect in dimer systems by proving the condensation of the system's bulk eigenmodes at one of the edges of the system; (iii) to show the topological origin of the non-Hermitian skin effect for dimer systems and (iv) to prove localisation of the interface modes between two dimer structures with non-Hermitian gauge potentials of opposite signs based on new estimates of the decay of the entries of the eigenvectors of block matrices with mirrored blocks.

math-ph↗

Text-guided Eyeglasses Manipulation with Spatial Constraints

Virtual try-on of eyeglasses involves placing eyeglasses of different shapes and styles onto a face image without physically trying them on. While existing methods have shown impressive results, the variety of eyeglasses styles is limited and the interactions are not always intuitive or efficient. To address these limitations, we propose a Text-guided Eyeglasses Manipulation method that allows for control of the eyeglasses shape and style based on a binary mask and text, respectively. Specifically, we introduce a mask encoder to extract mask conditions and a modulation module that enables simultaneous injection of text and mask conditions. This design allows for fine-grained control of the eyeglasses' appearance based on both textual descriptions and spatial constraints. Our approach includes a disentangled mapper and a decoupling strategy that preserves irrelevant areas, resulting in better local editing. We employ a two-stage training scheme to handle the different convergence speeds of the various modality conditions, successfully controlling both the shape and style of eyeglasses. Extensive comparison experiments and ablation analyses demonstrate the effectiveness of our approach in achieving diverse eyeglasses styles while preserving irrelevant areas.

cs.CV↗

Generating Reliable Pixel-Level Labels for Source Free Domain Adaptation

This work addresses the challenging domain adaptation setting in which knowledge from the labelled source domain dataset is available only from the pretrained black-box segmentation model. The pretrained model's predictions for the target domain images are noisy because of the distributional differences between the source domain data and the target domain data. Since the model's predictions serve as pseudo labels during self-training, the noise in the predictions impose an upper bound on model performance. Therefore, we propose a simple yet novel image translation workflow, ReGEN, to address this problem. ReGEN comprises an image-to-image translation network and a segmentation network. Our workflow generates target-like images using the noisy predictions from the original target domain images. These target-like images are semantically consistent with the noisy model predictions and therefore can be used to train the segmentation network. In addition to being semantically consistent with the predictions from the original target domain images, the generated target-like images are also stylistically similar to the target domain images. This allows us to leverage the stylistic differences between the target-like images and the target domain image as an additional source of supervision while training the segmentation model. We evaluate our model with two benchmark domain adaptation settings and demonstrate that our approach performs favourably relative to recent state-of-the-art work. The source code will be made available.

cs.CV↗

UTSGAN: Unseen Transition Suss GAN for Transition-Aware Image-to-image Translation

In the field of Image-to-Image (I2I) translation, ensuring consistency between input images and their translated results is a key requirement for producing high-quality and desirable outputs. Previous I2I methods have relied on result consistency, which enforces consistency between the translated results and the ground truth output, to achieve this goal. However, result consistency is limited in its ability to handle complex and unseen attribute changes in translation tasks. To address this issue, we introduce a transition-aware approach to I2I translation, where the data translation mapping is explicitly parameterized with a transition variable, allowing for the modelling of unobserved translations triggered by unseen transitions. Furthermore, we propose the use of transition consistency, defined on the transition variable, to enable regularization of consistency on unobserved translations, which is omitted in previous works. Based on these insights, we present Unseen Transition Suss GAN (UTSGAN), a generative framework that constructs a manifold for the transition with a stochastic transition encoder and coherently regularizes and generalizes result consistency and transition consistency on both training and unobserved translations with tailor-designed constraints. Extensive experiments on four different I2I tasks performed on five different datasets demonstrate the efficacy of our proposed UTSGAN in performing consistent translations.

cs.CV↗

An Operator Theory for Analyzing the Resolution of Multi-illumination Imaging Modalities

By introducing a new operator theory, we provide a unified mathematical theory for general source resolution in the multi-illumination imaging problem. Our main idea is to transform multi-illumination imaging into single-snapshot imaging with a new imaging kernel that depends on both the illumination patterns and the point spread function of the imaging system. We thus prove that the resolution of multi-illumination imaging is approximately determined by the essential cutoff frequency of the new imaging kernel, which is roughly limited by the sum of the cutoff frequency of the point spread function and the maximum essential frequency in the illumination patterns. Our theory provides a unified way to estimate the resolution of various existing super-resolution modalities and results in the same estimates as those obtained in experiments. In addition, based on the reformulation of the multi-illumination imaging problem, we also estimate the resolution limits for resolving both complex and positive sources by sparsity-based approaches. We show that the resolution of multi-illumination imaging is approximately determined by the new imaging kernel from our operator theory and better resolution can be realized by sparsity-promoting techniques in practice but only for resolving very sparse sources. This explains experimentally observed phenomena in some sparsity-based super-resolution modalities.

eess.IV↗

Super-resolution of positive near-colliding point sources

In this paper, we analyze the capacity of super-resolution of one-dimensional positive sources. In particular, we consider the same setting as in [arXiv:1904.09186v2 [math.NA]] and generalize the results there to the case of super-resolving positive sources. To be more specific, we consider resolving $d$ positive point sources with $p \leqslant d$ nodes closely spaced and forming a cluster, while the rest of the nodes are well separated. Similarly to [arXiv:1904.09186v2 [math.NA]], our results show that when the noise level $ε\lesssim \mathrm{SRF}^{-2 p+1}$, where $\mathrm{SRF}=(ΩΔ)^{-1}$ with $Ω$ being the cutoff frequency and $Δ$ the minimal separation between the nodes, the minimax error rate for reconstructing the cluster nodes is of order $\frac{1}Ω \mathrm{SRF}^{2 p-2} ε$, while for recovering the corresponding amplitudes $\left\{a_j\right\}$ the rate is of order $\mathrm{SRF}^{2 p-1} ε$. For the non-cluster nodes, the corresponding minimax rates for the recovery of nodes and amplitudes are of order $\fracεΩ$ and $ε$, respectively. Our numerical experiments show that the Matrix Pencil method achieves the above optimal bounds when resolving the positive sources.

eess.IV↗

SPG-VTON: Semantic Prediction Guidance for Multi-pose Virtual Try-on

Image-based virtual try-on is challenging in fitting a target in-shop clothes into a reference person under diverse human poses. Previous works focus on preserving clothing details ( e.g., texture, logos, patterns ) when transferring desired clothes onto a target person under a fixed pose. However, the performances of existing methods significantly dropped when extending existing methods to multi-pose virtual try-on. In this paper, we propose an end-to-end Semantic Prediction Guidance multi-pose Virtual Try-On Network (SPG-VTON), which could fit the desired clothing into a reference person under arbitrary poses. Concretely, SPG-VTON is composed of three sub-modules. First, a Semantic Prediction Module (SPM) generates the desired semantic map. The predicted semantic map provides more abundant guidance to locate the desired clothes region and produce a coarse try-on image. Second, a Clothes Warping Module (CWM) warps in-shop clothes to the desired shape according to the predicted semantic map and the desired pose. Specifically, we introduce a conductible cycle consistency loss to alleviate the misalignment in the clothes warping process. Third, a Try-on Synthesis Module (TSM) combines the coarse result and the warped clothes to generate the final virtual try-on image, preserving details of the desired clothes and under the desired pose. Besides, we introduce a face identity loss to refine the facial appearance and maintain the identity of the final virtual try-on result at the same time. We evaluate the proposed method on the most massive multi-pose dataset (MPV) and the DeepFashion dataset. The qualitative and quantitative experiments show that SPG-VTON is superior to the state-of-the-art methods and is robust to the data noise, including background and accessory changes, i.e., hats and handbags, showing good scalability to the real-world scenario.

cs.CV↗

Meta Knowledge Condensation for Federated Learning

Existing federated learning paradigms usually extensively exchange distributed models at a central solver to achieve a more powerful model. However, this would incur severe communication burden between a server and multiple clients especially when data distributions are heterogeneous. As a result, current federated learning methods often require a large number of communication rounds in training. Unlike existing paradigms, we introduce an alternative perspective to significantly decrease the communication cost in federate learning. In this work, we first introduce a meta knowledge representation method that extracts meta knowledge from distributed clients. The extracted meta knowledge encodes essential information that can be used to improve the current model. As the training progresses, the contributions of training samples to a federated model also vary. Thus, we introduce a dynamic weight assignment mechanism that enables samples to contribute adaptively to the current model update. Then, informative meta knowledge from all active clients is sent to the server for model update. Training a model on the combined meta knowledge without exposing original data among different clients can significantly mitigate the heterogeneity issues. Moreover, to further ameliorate data heterogeneity, we also exchange meta knowledge among clients as conditional initialization for local meta knowledge extraction. Extensive experiments demonstrate the effectiveness and efficiency of our proposed method. Remarkably, our method outperforms the state-of-the-art by a large margin (from $74.07\%$ to $92.95\%$) on MNIST with a restricted communication budget (i.e. 10 rounds).

cs.LG↗

Improved Multi-Dimensional Bee Colony Algorithm for Airport Freight Station Scheduling

Due to the rapid increase of air cargo and postal transport volume, an efficient automated multi-dimensional warehouse with elevating transfer vehicles (ETVs) should be established and an effective scheduling strategy should be designed for improving the cargo handling efficiency. In this paper, artificial bee colony algorithm, which possesses strong global optimization ability and fewer parameters, is firstly introduced to simultaneously optimize the route of ETV and the assignment of entrances and exits. Moreover, for further improve the optimization performance of ABC, novel full-dimensional search strategy with parallelization, and random multi-dimensional search strategy are incorporated in the framework of ABC to improve the diversity of the population and the convergence speed respectively. Our proposed algorithms are evaluated on several benchmark functions, and then applied to solve the combinatorial optimization problem with multitask, multiple entrances and exits in air cargo terminal. The simulations show that the proposed algorithms can achieve much more desired performance than the traditional artificial bee colony algorithm at balancing the exploitation and exploration abilities.

math.OC↗