arXiv ScienceSearch

arXiv subjects

Xianfeng Zhao

Publications and source records attributed to Xianfeng Zhao.

At least 19 recordsLinked to original sources

The Douglas question for functions of the form $\bar{z}+h$ with $h$ in the disk algebra

We prove that the Toeplitz operator $T_{\bar{z}+h}$ is invertible on the Bergman space provided that $|\bar{z}+h|$ is bounded below by some positive constant on the open unit disk, where $h$ is in the disk algebra. This provides a general class of harmonic functions for which the answer to the Douglas question on the Bergman space is affirmative.

math.FA

The Douglas question on the Bergman and Fock spaces

Let $\mu$ be a positive Borel measure and $T_\mu$ be the bounded Toeplitz operator induced by $\mu$ on the Bergman or Fock space. In this paper, we mainly investigate the invertibility of the Toeplitz operator $T_\mu$ and the Douglas question on the Bergman and Fock spaces. In the Bergman-space setting, we obtain several necessary and sufficient conditions for the invertibility of $T_\mu$ in terms of the Berezin transform of $\mu$ and the reverse Carleson condition in two classical cases: (1) $\mu$ is absolutely continuous with respect to the normalized area measure on the open unit disk $\mathbb D$; (2) $\mu$ is the pull-back measure of the normalized area measure under an analytic self-mapping of $\mathbb D$. Nonetheless, we show that there exists a Carleson measure for the Bergman space such that its Berezin transform is bounded below but the corresponding Toeplitz operator is not invertible. On the Fock space, we show that $T_\mu$ is invertible if and only if $\mu$ is a reverse Carleson measure, but the invertibility of $T_\mu$ is not completely determined by the invertibility of the Berezin transform of $\mu$. These suggest that the answers to the Douglas question for Toeplitz operators induced by positive measures on the Bergman and Fock spaces are both negative in general cases.

math.FA

ZeroGen: Zero-shot Multimodal Controllable Text Generation with Multiple Oracles

Automatically generating textual content with desired attributes is an ambitious task that people have pursued long. Existing works have made a series of progress in incorporating unimodal controls into language models (LMs), whereas how to generate controllable sentences with multimodal signals and high efficiency remains an open question. To tackle the puzzle, we propose a new paradigm of zero-shot controllable text generation with multimodal signals (\textsc{ZeroGen}). Specifically, \textsc{ZeroGen} leverages controls of text and image successively from token-level to sentence-level and maps them into a unified probability space at decoding, which customizes the LM outputs by weighted addition without extra training. To achieve better inter-modal trade-offs, we further introduce an effective dynamic weighting mechanism to regulate all control weights. Moreover, we conduct substantial experiments to probe the relationship of being in-depth or in-width between signals from distinct modalities. Encouraging empirical results on three downstream tasks show that \textsc{ZeroGen} not only outperforms its counterparts on captioning tasks by a large margin but also shows great potential in multimodal news generation with a higher degree of control. Our code will be released at https://github.com/ImKeTT/ZeroGen.

cs.CL

Errorless Robust JPEG Steganography Using Steganographic Polar Codes

Recently, a robust steganographic algorithm that achieves errorless robustness against JPEG recompression is proposed. The method evaluates the behavior of DCT coefficients after recompression using the local JPEG encoder to select robust coefficients and sets the other coefficients as wet cost. Combining the lattice embedding scheme, the method is errorless by construction. However, the authors only concern with the success rate under theoretical embedding, while the success rate of the implementation with practical steganographic codes is not verified. In this letter, we implement the method with two steganographic codes, i.e., steganographic polar code and syndrome-trellis code. By analyzing the possibility of success embedding of two steganographic codes under wet paper embedding, we discover that steganographic polar code achieves success embedding with a larger number of wet coefficients compared with syndrome-trellis code, which makes steganographic polar code more suitable under the errorless robust embedding paradigm. The experimental results show that the combination of steganographic polar code and errorless robust embedding achieves a higher success rate compared with the implementation with syndrome-trellis code under close security performance.

cs.CR

Toeplitz operators on $\mathcal L^p$-spaces of a tree

Let $T$ be a rooted, countable infinite tree without terminal vertices. In the present paper, we characterize the spectra, self-adjointness and positivity of Toeplitz operators on the spaces of $p$-summable functions on $T$. Moreover, we obtain a necessary and sufficient condition for Toeplitz operators to have finite rank on such function spaces.

math.FA

Improving Robustness of TCM-based Robust Steganography with Variable Robustness

Recent study has found out that after multiple times of recompression, the DCT coefficients of JPEG image can form an embedding domain that is robust to recompression, which is called transport channel matching (TCM) method. Because the cost function of the adaptive steganography does not consider the impact of modification on the robustness, the modified DCT coefficients of the stego image after TCM will change after recompression. To reduce the number of changed coefficients after recompression, this paper proposes a robust steganography algorithm which dynamically updates the robustness cost of every DCT coefficient. The robustness cost proposed is calculated by testing whether the modified DCT coefficient can resist recompression in every step of STC embedding process. By adding robustness cost to the distortion cost and using the framework of STC embedding algorithm to embed the message, the stego images have good performance both in robustness and security. The experimental results show that the proposed algorithm can significantly enhance the robustness of stego images, and the embedded messages could be extracted correctly at almost all cases when recompressing with a lower quality factor and recompression process is known to the user of proposed algorithm.

cs.CR

Vision Transformer Based Video Hashing Retrieval for Tracing the Source of Fake Videos

In recent years, the spread of fake videos has brought great influence on individuals and even countries. It is important to provide robust and reliable results for fake videos. The results of conventional detection methods are not reliable and not robust for unseen videos. Another alternative and more effective way is to find the original video of the fake video. For example, fake videos from the Russia-Ukraine war and the Hong Kong law revision storm are refuted by finding the original video. We use an improved retrieval method to find the original video, named ViTHash. Specifically, tracing the source of fake videos requires finding the unique one, which is difficult when there are only small differences in the original videos. To solve the above problems, we designed a novel loss Hash Triplet Loss. In addition, we designed a tool called Localizator to compare the difference between the original traced video and the fake video. We have done extensive experiments on FaceForensics++, Celeb-DF and DeepFakeDetection, and we also have done additional experiments on our built three datasets: DAVIS2016-TL (video inpainting), VSTL (video splicing) and DFTL (similar videos). Experiments have shown that our performance is better than state-of-the-art methods, especially in cross-dataset mode. Experiments also demonstrated that ViTHash is effective in various forgery detection: video inpainting, video splicing and deepfakes. Our code and datasets have been released on GitHub: \url{https://github.com/lajlksdf/vtl}.

cs.CV

FMFCC-A: A Challenging Mandarin Dataset for Synthetic Speech Detection

As increasing development of text-to-speech (TTS) and voice conversion (VC) technologies, the detection of synthetic speech has been suffered dramatically. In order to promote the development of synthetic speech detection model against Mandarin TTS and VC technologies, we have constructed a challenging Mandarin dataset and organized the accompanying audio track of the first fake media forensic challenge of China Society of Image and Graphics (FMFCC-A). The FMFCC-A dataset is by far the largest publicly-available Mandarin dataset for synthetic speech detection, which contains 40,000 synthesized Mandarin utterances that generated by 11 Mandarin TTS systems and two Mandarin VC systems, and 10,000 genuine Mandarin utterances collected from 58 speakers. The FMFCC-A dataset is divided into the training, development and evaluation sets, which are used for the research of detection of synthesized Mandarin speech under various previously unknown speech synthesis systems or audio post-processing operations. In addition to describing the construction of the FMFCC-A dataset, we provide a detailed analysis of two baseline methods and the top-performing submissions from the FMFCC-A, which illustrates the usefulness and challenge of FMFCC-A dataset. We hope that the FMFCC-A dataset can fill the gap of lack of Mandarin datasets for synthetic speech detection.

cs.SD

MediumVC: Any-to-any voice conversion using synthetic specific-speaker speeches as intermedium features

To realize any-to-any (A2A) voice conversion (VC), most methods are to perform symmetric self-supervised reconstruction tasks (Xi to Xi), which usually results in inefficient performances due to inadequate feature decoupling, especially for unseen speakers. We propose a two-stage reconstruction task (Xi to Yi to Xi) using synthetic specific-speaker speeches as intermedium features, where A2A VC is divided into two stages: any-to-one (A2O) and one-to-Any (O2A). In the A2O stage, we propose a new A2O method: SingleVC, by employing a noval data augment strategy(pitch-shifted and duration-remained, PSDR) to accomplish Xi to Yi. In the O2A stage, MediumVC is proposed based on pre-trained SingleVC to conduct Yi to Xi. Through such asymmetrical reconstruction tasks (Xi to Yi in SingleVC and Yi to Xi in MediumVC), the models are to capture robust disentangled features purposefully. Experiments indicate MediumVC can enhance the similarity of converted speeches while maintaining a high degree of naturalness.

eess.AS

Detection of Deepfake Videos Using Long Distance Attention

With the rapid progress of deepfake techniques in recent years, facial video forgery can generate highly deceptive video contents and bring severe security threats. And detection of such forgery videos is much more urgent and challenging. Most existing detection methods treat the problem as a vanilla binary classification problem. In this paper, the problem is treated as a special fine-grained classification problem since the differences between fake and real faces are very subtle. It is observed that most existing face forgery methods left some common artifacts in the spatial domain and time domain, including generative defects in the spatial domain and inter-frame inconsistencies in the time domain. And a spatial-temporal model is proposed which has two components for capturing spatial and temporal forgery traces in global perspective respectively. The two components are designed using a novel long distance attention mechanism. The one component of the spatial domain is used to capture artifacts in a single frame, and the other component of the time domain is used to capture artifacts in consecutive frames. They generate attention maps in the form of patches. The attention method has a broader vision which contributes to better assembling global information and extracting local statistic information. Finally, the attention maps are used to guide the network to focus on pivotal parts of the face, just like other fine-grained classification methods. The experimental results on different public datasets demonstrate that the proposed method achieves the state-of-the-art performance, and the proposed long distance attention method can effectively capture pivotal parts for face forgery.

cs.CV

Toeplitz algebras over Fock and Bergman spaces

In this paper, we study Toeplitz algebras generated by certain class of Toeplitz operators on the $p$-Fock space and the $p$-Bergman space with $1<p<\infty$. Let BUC($\mathbb C^n$) and BUC($\mathbb B_n$) denote the collections of bounded uniformly continuous functions on $\mathbb C^n$ and $\mathbb B_n$ (the unit ball in $\mathbb C^n$), respectively. On the $p$-Fock space, we show that the Toeplitz algebra which has a translation invariant closed subalgebra of BUC($\mathbb C^n$) as its set of symbols is linearly generated by Toeplitz operators with the same space of symbols. This answers a question recently posed by Fulsche \cite{Robert}. On the $p$-Bergman space, we study Toeplitz algebras with symbols in some translation invariant closed subalgebras of BUC($\mathbb B_n)$. In particular, we obtain that the Toeplitz algebra generated by all Toeplitz operators with symbols in BUC($\mathbb B_n$) is equal to the closed linear space generated by Toeplitz operators with such symbols. This generalizes the corresponding result for the case of $p=2$ obtained by Xia \cite{Xia2015}.

math.FA

Hyponormal dual Toeplitz operators on the orthogonal complement of the Harmonic Bergman space

In this paper, we mainly study the hyponormality of dual Toeplitz operators on the orthogonal complement of the harmonic Bergman space. First we show that the dual Toeplitz operator with bounded symbol is hyponormal if and only if it is normal. Then we obtain a necessary and sufficient condition for the dual Toeplitz operator $S_\varphi$ with the symbol $\varphi(z) = az^{n_1}\overline{z}^{m_1} + bz^{n_2} \overline{z}^{m_2}$, $(n_1,n_2,m_1,m_2\in \mathbb {N}$ and $a,b \in \mathbb{C})$ to be hyponormal. Finally, we show that the rank of the commutator of two dual Toeplitz operators must be an even number if the commutator has a finite rank.

math.FA

Essentially commuting dual truncated Toeplitz operators

In this paper, we completely characterize when two dual truncated Toeplitz operators are essentially commuting and when the semicommutator of two dual truncated Toeplitz operators is compact. Our main idea is to study dual truncated Toeplitz operators via Hankel operators, Toeplitz operators and function algebras.

math.FA

The spectral picture of Bergman Toeplitz operators with harmonic polynomial symbols

In this paper, it is shown that some new phenomenon related to the spectra of Toeplitz operators with bounded harmonic symbols on the Bergman space. On the one hand, we prove that the spectrum of the Toeplitz operator with symbol ${\bar{z}+p}$ is always connected for every polynomial $p$ with degree less than $3$. On the other hand, we show that for each integer $k$ greater than $2$, there exists a polynomial $p$ of degree $k$ such that the spectrum of the Toeplitz operator with symbol ${\bar{z}+p}$ has at least one isolated point but has at most finitely many isolated points. Then these results are applied to obtain a new class of non-hyponormal Toeplitz operators with bounded harmonic symbols on the Bergman space for which Weyl's theorem holds.

math.FA

IStego100K: Large-scale Image Steganalysis Dataset

In order to promote the rapid development of image steganalysis technology, in this paper, we construct and release a multivariable large-scale image steganalysis dataset called IStego100K. It contains 208,104 images with the same size of 1024*1024. Among them, 200,000 images (100,000 cover-stego image pairs) are divided as the training set and the remaining 8,104 as testing set. In addition, we hope that IStego100K can help researchers further explore the development of universal image steganalysis algorithms, so we try to reduce limits on the images in IStego100K. For each image in IStego100K, the quality factors is randomly set in the range of 75-95, the steganographic algorithm is randomly selected from three well-known steganographic algorithms, which are J-uniward, nsF5 and UERD, and the embedding rate is also randomly set to be a value of 0.1-0.4. In addition, considering the possible mismatch between training samples and test samples in real environment, we add a test set (DS-Test) whose source of samples are different from the training set. We hope that this test set can help to evaluate the robustness of steganalysis algorithms. We tested the performance of some latest steganalysis algorithms on IStego100K, with specific results and analysis details in the experimental part. We hope that the IStego100K dataset will further promote the development of universal image steganalysis technology. The description of IStego100K and instructions for use can be found at https://github.com/YangzlTHU/IStego100K

cs.CR

Adversarial Learning for Image Forensics Deep Matching with Atrous Convolution

Constrained image splicing detection and localization (CISDL) is a newly proposed challenging task for image forensics, which investigates two input suspected images and identifies whether one image has suspected regions pasted from the other. In this paper, we propose a novel adversarial learning framework to train the deep matching network for CISDL. Our framework mainly consists of three building blocks: 1) the deep matching network based on atrous convolution (DMAC) aims to generate two high-quality candidate masks which indicate the suspected regions of the two input images, 2) the detection network is designed to rectify inconsistencies between the two corresponding candidate masks, 3) the discriminative network drives the DMAC network to produce masks that are hard to distinguish from ground-truth ones. In DMAC, atrous convolution is adopted to extract features with rich spatial information, the correlation layer based on the skip architecture is proposed to capture hierarchical features, and atrous spatial pyramid pooling is constructed to localize tampered regions at multiple scales. The detection network and the discriminative network act as the losses with auxiliary parameters to supervise the training of DMAC in an adversarial way. Extensive experiments, conducted on 21 generated testing sets and two public datasets, demonstrate the effectiveness of the proposed framework and the superior performance of DMAC.

cs.CV

Adaptive Spatial Steganography Based on Probability-Controlled Adversarial Examples

Explanation from Sai Ma: The experiments in this paper are conducted on Caffe framework. In Caffe, there is an API to directly set the gradient in Matlab. I wrongly use it to control the 'probability', in fact, I modify the gradient directly. The misusage of API leads to wrong experiment results, and wrong theoretical analysis. Apologize to readers who have read this paper. We have submitted a correct version of this paper to Multimedia Tools and Applications and it is under revision. Thanks to Dr. Patrick Bas, who is the Associate Editor of TIFS and the anonymous reviewers of this paper. Thanks to Tingting Song from Sun Yat-sen University. We discussed some problems of this paper. Her advice helps me to improve the submitted paper to Multimedia Tools and Applications.

cs.MM

Weakening the Detecting Capability of CNN-based Steganalysis

Recently, the application of deep learning in steganalysis has drawn many researchers' attention. Most of the proposed steganalytic deep learning models are derived from neural networks applied in computer vision. These kinds of neural networks have distinguished performance. However, all these kinds of back-propagation based neural networks may be cheated by forging input named the adversarial example. In this paper we propose a method to generate steganographic adversarial example in order to enhance the steganographic security of existing algorithms. These adversarial examples can increase the detection error of steganalytic CNN. The experiments prove the effectiveness of the proposed method.

cs.MM