arXiv Science⌕ Search

arXiv subjects

Chen Cheng

Publications and source records attributed to Chen Cheng.

At least 73 records · Page 4Linked to original sources

Exploring the Robustness of In-Context Learning with Noisy Labels

Recently, the mysterious In-Context Learning (ICL) ability exhibited by Transformer architectures, especially in large language models (LLMs), has sparked significant research interest. However, the resilience of Transformers' in-context learning capabilities in the presence of noisy samples, prevalent in both training corpora and prompt demonstrations, remains underexplored. In this paper, inspired by prior research that studies ICL ability using simple function classes, we take a closer look at this problem by investigating the robustness of Transformers against noisy labels. Specifically, we first conduct a thorough evaluation and analysis of the robustness of Transformers against noisy labels during in-context learning and show that they exhibit notable resilience against diverse types of noise in demonstration labels. Furthermore, we delve deeper into this problem by exploring whether introducing noise into the training set, akin to a form of data augmentation, enhances such robustness during inference, and find that such noise can indeed improve the robustness of ICL. Overall, our fruitful analysis and findings provide a comprehensive understanding of the resilience of Transformer models against label noises during ICL and provide valuable insights into the research on Transformers in natural language processing. Our code is available at https://github.com/InezYu0928/in-context-learning.

cs.CL↗

Searching for Two-Neutrino and Neutrinoless Double Beta Decay of $^{134}$Xe with the PandaX-4T Experiment

$^{134}$Xe is a candidate isotope for neutrinoless double beta decay~($0νββ$) search. In addition, the two-neutrino case ($2νββ$) allowed by the Standard Model of particle physics has not yet been observed. Utilizing the 10.4% of $^{134}$Xe in the natural xenon in the PandaX-4T detector and its first 94.9-day exposure, we have established the most stringent constraints on $2νββ$ and $0νββ$ of $^{134}$Xe half-lives, with limits of $2.8\times10^{22}$ yr and $3.0\times10^{23}$ yr at 90% confidence level, respectively. The $2νββ$ ($0νββ$) limit surpasses the previously reported best result by a factor of 32 (2.7), highlighting the potential of large monolithic natural xenon detectors.

nucl-ex↗

Detecting Neutrinos from Supernova Bursts in PandaX-4T

Neutrinos from core-collapse supernovae are essential for the understanding of neutrino physics and stellar evolution. The dual-phase xenon dark matter detectors can provide a way to track explosions of galactic supernovae by detecting neutrinos through coherent elastic neutrino-nucleus scatterings. In this study, a variation of progenitor masses as well as explosion models are assumed to predict the neutrino fluxes and spectra, which result in the number of expected neutrino events ranging from 6.6 to 13.7 at a distance of 10 kpc over a 10-second duration with negligible backgrounds at PandaX-4T. Two specialized triggering alarms for monitoring supernova burst neutrinos are built. The efficiency of detecting supernova explosions at various distances in the Milky Way is estimated. These alarms will be implemented in the real-time supernova monitoring system at PandaX-4T in the near future, providing the astronomical communities with supernova early warnings.

hep-ex↗

Search for Dark-Matter-Nucleon Interactions with a Dark Mediator in PandaX-4T

We report results of a search for dark-matter-nucleon interactions via a dark mediator using optimized low-energy data from the PandaX-4T liquid xenon experiment. With the ionization-signal-only data and utilizing the Migdal effect, we set the most stringent limits on the cross section for dark matter masses ranging from 30~$\rm{MeV/c^2}$ to 2~$\rm{GeV/c^2}$. Under the assumption that the dark mediator is a dark photon that decays into scalar dark matter pairs in the early Universe, we rule out significant parameter space of such thermal relic dark-matter model.

hep-ex↗

FaceChain: A Playground for Human-centric Artificial Intelligence Generated Content

Recent advancement in personalized image generation have unveiled the intriguing capability of pre-trained text-to-image models on learning identity information from a collection of portrait images. However, existing solutions are vulnerable in producing truthful details, and usually suffer from several defects such as (i) The generated face exhibit its own unique characteristics, \ie facial shape and facial feature positioning may not resemble key characteristics of the input, and (ii) The synthesized face may contain warped, blurred or corrupted regions. In this paper, we present FaceChain, a personalized portrait generation framework that combines a series of customized image-generation model and a rich set of face-related perceptual understanding models (\eg, face detection, deep face embedding extraction, and facial attribute recognition), to tackle aforementioned challenges and to generate truthful personalized portraits, with only a handful of portrait images as input. Concretely, we inject several SOTA face models into the generation procedure, achieving a more efficient label-tagging, data-processing, and model post-processing compared to previous solutions, such as DreamBooth ~\cite{ruiz2023dreambooth} , InstantBooth ~\cite{shi2023instantbooth} , or other LoRA-only approaches ~\cite{hu2021lora} . Besides, based on FaceChain, we further develop several applications to build a broader playground for better showing its value, including virtual try on and 2D talking head. We hope it can grow to serve the burgeoning needs from the communities. Note that this is an ongoing work that will be consistently refined and improved upon. FaceChain is open-sourced under Apache-2.0 license at \url{https://github.com/modelscope/facechain}.

cs.CV↗

CUCL: Codebook for Unsupervised Continual Learning

The focus of this study is on Unsupervised Continual Learning (UCL), as it presents an alternative to Supervised Continual Learning which needs high-quality manual labeled data. The experiments under the UCL paradigm indicate a phenomenon where the results on the first few tasks are suboptimal. This phenomenon can render the model inappropriate for practical applications. To address this issue, after analyzing the phenomenon and identifying the lack of diversity as a vital factor, we propose a method named Codebook for Unsupervised Continual Learning (CUCL) which promotes the model to learn discriminative features to complete the class boundary. Specifically, we first introduce a Product Quantization to inject diversity into the representation and apply a cross quantized contrastive loss between the original representation and the quantized one to capture discriminative information. Then, based on the quantizer, we propose an effective Codebook Rehearsal to address catastrophic forgetting. This study involves conducting extensive experiments on CIFAR100, TinyImageNet, and MiniImageNet benchmark datasets. Our method significantly boosts the performances of supervised and unsupervised methods. For instance, on TinyImageNet, our method led to a relative improvement of 12.76% and 7% when compared with Simsiam and BYOL, respectively.

cs.CV↗

Multi-dimensional vibration sensing and simultaneous self-homodyne optical transmission of single wavelength net 5.36 Tb/s signal using telecom 7-core fiber

We present a high-capacity self-homodyne optical transmission system that enables simultaneously multidimensional vibration sensing based on a weakly-coupled 7-core fiber. To our knowledge, we demonstrate for the first-time detection of fiber vibration direction along with strength, frequency, and location of the vibration source, while transmitting in the meantime single-carrier 16 QAM signal reaching a net date rate of 5.36 Tb/s over 41.4 km of telecom 7-core fiber.

physics.optics↗

Investigating Berezinskii-Kosterlitz-Thouless phase transitions in Kagome spin ice by quantifying Monte Carlo process: Distribution of Hamming distances

We reinvestigate the phase transitions of the Ising model on the Kagome lattice with antiferromagnetic nearest-neighbor and ferromagnetic next-nearest-neighbor interactions, which has a six-state-clock spin ice ground state and two consecutive Berezinskii-Kosterlitz-Thouless (BKT) phase transitions. Employing the classical Monte Carlo (MC) simulations, the phases are characterized by the magnetic order parameter, and the critical temperatures are obtained by the finite-size scaling of related physical quantities. Moreover, we attempt to gain general information on the phase transitions from the MC process instead of MC results and successfully extract the correct transition points with surprisingly high accuracy. Specifically, we focus on the selected data set of uncorrelated MC configurations and quantify the MC process using the distribution of two-configuration Hamming distances in this small data collection. This distribution is more than a quantity that features different behaviors in different phases but also nicely supports the same BKT scaling form as the order parameter, from which we successfully determine the two BKT transition points with surprisingly high accuracy. We also discuss the connection between the phase transitions and the intrinsic dimension extracted from the Hamming distances, which is widely used in the growing field of machine learning and is reported to be able to detect critical points. Our findings provide a new understanding of the spin ice transitions in the Kagome lattice and can hopefully be used similarly to identify transitions in the quantum system on the same lattice with strong frustrations.

cond-mat.stat-mech↗

An explicit evolution from Néel to striped antiferromagnetic states in the spin-1/2 $J_{1}$-$J_{2}$ Heisenberg model on the square lattice

The frustrated spin-$1/2$ $J_1-J_2$ Heisenberg model on the square lattice has been extensively studied since 1988 because of its close relationship to the high-temperature superconductivity in cuprates and more importantly involved novel phase of matter in its own right, namely, quantum spin liquid (QSL), one of hot topics in condensed matter physics in recent years. However, the phase diagram of the model, particularly in the maximally frustrated regime $J_2/J_1 \sim 0.5$, is quite controversial, and more seriously the nature of the QSL is not clear at all. Here we provide a pattern picture, on one hand, to show explicitly how the system evolves from the Néel antiferromagnetic (AFM) state at small $J_2$ to the striped AFM one at large $J_2$; on the other hand, to uncover the nature of the QSL if it exists in the intermediate $J_2$ coupling regime. For simplicity, we show our results by taking the square lattice $L=L_x \times L_y$ with size $L_x=L_y=4$ here and periodic boundary condition is considered, and furthermore, exact diagonalization is employed to confirm the correctness of our picture. Our results indicate that the highly frustration regime is characterized by diagonal two-domain, while the Néel AFM state has a diagonal single-domain and the striped AFM state shows itself as a diagonal four-domain, namely, completely diagonal antiferromagnetic order, in the present case. Increasing the system size, the number of the diagonal domains increases correspondingly, but the diagonal single-domain for the Néel AFM state and the diagonal $L_{x(y)}$-domain for the striped AFM state remain unchanged. Our results shed light on the understanding of the QSL.

cond-mat.str-el↗

ModelScope-Agent: Building Your Customizable Agent System with Open-source Large Language Models

Large language models (LLMs) have recently demonstrated remarkable capabilities to comprehend human intentions, engage in reasoning, and design planning-like behavior. To further unleash the power of LLMs to accomplish complex tasks, there is a growing trend to build agent framework that equips LLMs, such as ChatGPT, with tool-use abilities to connect with massive external APIs. In this work, we introduce ModelScope-Agent, a general and customizable agent framework for real-world applications, based on open-source LLMs as controllers. It provides a user-friendly system library, with customizable engine design to support model training on multiple open-source LLMs, while also enabling seamless integration with both model APIs and common APIs in a unified way. To equip the LLMs with tool-use abilities, a comprehensive framework has been proposed spanning over tool-use data collection, tool retrieval, tool registration, memory control, customized model training, and evaluation for practical real-world applications. Finally, we showcase ModelScopeGPT, a real-world intelligent assistant of ModelScope Community based on the ModelScope-Agent framework, which is able to connect open-source LLMs with more than 1000 public AI models and localized community knowledge in ModelScope. The ModelScope-Agent library\footnote{https://github.com/modelscope/modelscope-agent} and online demo\footnote{https://modelscope.cn/studios/damo/ModelScopeGPT/summary} are now publicly available.

cs.CL↗

ALens: An Adaptive Domain-Oriented Abstract Writing Training Tool for Novice Researchers

The significance of novice researchers acquiring proficiency in writing abstracts has been extensively documented in the field of higher education, where they often encounter challenges in this process. Traditionally, students have been advised to enroll in writing training courses as a means to develop their abstract writing skills. Nevertheless, this approach frequently falls short in providing students with personalized and adaptable feedback on their abstract writing. To address this gap, we initially conducted a formative study to ascertain the user requirements for an abstract writing training tool. Subsequently, we proposed a domain-specific abstract writing training tool called ALens, which employs rhetorical structure parsing to identify key concepts, evaluates abstract drafts based on linguistic features, and employs visualization techniques to analyze the writing patterns of exemplary abstracts. A comparative user study involving an alternative abstract writing training tool has been conducted to demonstrate the efficacy of our approach.

cs.HC↗

Search for light dark matter from atmosphere in PandaX-4T

We report a search for light dark matter produced through the cascading decay of $η$ mesons, which are created as a result of inelastic collisions between cosmic rays and Earth's atmosphere. We introduce a new and general framework, publicly accessible, designed to address boosted dark matter specifically, with which a full and dedicated simulation including both elastic and quasi-elastic processes of Earth attenuation effect on the dark matter particles arriving at the detector is performed. In the PandaX-4T commissioning data of 0.63 tonne$\cdot$year exposure, no significant excess over background is observed. The first constraints on the interaction between light dark matter generated in the atmosphere and nucleus through a light scalar mediator are obtained. The lowest excluded cross-section is set at $5.9 \times 10^{-37}{\rm cm^2}$ for dark matter mass of $0.1$ MeV$/c^2$ and mediator mass of 300 MeV$/c^2$. The lowest upper limit of $η$ to dark matter decay branching ratio is $1.6 \times 10^{-7}$.

hep-ex↗

Collaboratively Learning Linear Models with Structured Missing Data

We study the problem of collaboratively learning least squares estimates for $m$ agents. Each agent observes a different subset of the features$\unicode{x2013}$e.g., containing data collected from sensors of varying resolution. Our goal is to determine how to coordinate the agents in order to produce the best estimator for each agent. We propose a distributed, semi-supervised algorithm Collab, consisting of three steps: local training, aggregation, and distribution. Our procedure does not require communicating the labeled data, making it communication efficient and useful in settings where the labeled data is inaccessible. Despite this handicap, our procedure is nearly asymptotically local minimax optimal$\unicode{x2013}$even among estimators allowed to communicate the labeled data such as imputation methods. We test our method on real and synthetic data.

stat.ML↗

Disorder in interacting quasi-one-dimensional systems: flat and dispersive bands

We investigate the superconductor-insulator transition (SIT) in disordered quasi-one dimensional systems using the density-matrix renormalization group method. Focusing on the case of an interacting spinful Hamiltonian at quarter-filling, we contrast the differences arising in the SIT when the parent non-interacting model features either flat or dispersive bands. Furthermore, by comparing disorder distributions that preserve or not SU(2)-symmetry, we unveil the critical disorder amplitude that triggers insulating behavior. While scaling analysis suggests the transition to be of a Berezinskii-Kosterlitz-Thouless type for all models (two lattices and two disorder types), only in the flat-band model with Zeeman-like disorder the critical disorder is nonvanishing. In this sense, the flat-band structure does strengthen superconductivity. For both flat and dispersive band models, i) in the presence of SU(2)-symmetric random chemical potentials, the disorder-induced transition is from superconductor to insulator of singlet pairs; ii) for the Zeeman-type disorder, the transition is from superconductor to insulator of unpaired fermions. In all cases, our numerical results suggest no intermediate disorder-driven metallic phase.

cond-mat.str-el↗

Parameter estimation from aggregate observations: A Wasserstein distance based sequential Monte Carlo sampler

In this work we study systems consisting of a group of moving particles. In such systems, often some important parameters are unknown and have to be estimated from observed data. Such parameter estimation problems can often be solved via a Bayesian inference framework. However in many practical problems, only data at the aggregate level is available and as a result the likelihood function is not available, which poses challenge for Bayesian methods. In particular, we consider the situation where the distributions of the particles are observed. We propose a Wasserstein distance based sequential Monte Carlo sampler to solve the problem: the Wasserstein distance is used to measure the similarity between the observed and the simulated particle distributions and the sequential Monte Carlo samplers is used to deal with the sequentially available observations. Two real-world examples are provided to demonstrate the performance of the proposed method.

stat.AP↗

A Search for Light Fermionic Dark Matter Absorption on Electrons in PandaX-4T

We report a search on a sub-MeV fermionic dark matter absorbed by electrons with an outgoing active neutrino using the 0.63 tonne-year exposure collected by PandaX-4T liquid xenon experiment. No significant signals are observed over the expected background. The data are interpreted into limits to the effective couplings between such dark matter and electrons. For axial-vector or vector interactions, our sensitivity is competitive in comparison to existing astrophysical bounds on the decay of such dark matter into photon final states. In particular, we present the first direct detection limits for an axial-vector (vector) interaction which are the strongest in the mass range from 25 to 45 (35 to 50) keV/c$^2$.

hep-ex↗

Many-body Localization in Clean Chains with Long-Range Interactions

The strong long-range interaction leads to localization in the closed quantum system without disorders. Employing the exact diagonalization method, the author numerically investigates thermalization and many-body localization in translational invariant quantum chains with finite Coulomb interactions. In the computational basis, excluding all trivial degeneracies, the interaction-induced localization is well demonstrated in aspects of level statistics, eigenstate expectation values, and the Anderson localization on graphs constructed of the many-body basis. The nature of localization for generic eigenstates is attributed to the quasi-disorder from the power-law interactions. However, due to the real-space symmetries, the long-time dynamics is dominated by the degenerate eigenstates and eventually reach homogeneity in real space. On the other hand, the entanglement entropy exhibits the size-dependence beyond the area law for the same reason, even deep in the localized state, indicating an incomplete localization in real space.

cond-mat.dis-nn↗

MCDIP-ADMM: Overcoming Overfitting in DIP-based CT reconstruction

This paper investigates the application of unsupervised learning methods for computed tomography (CT) reconstruction. To motivate our work, we review several existing priors, namely the truncated Gaussian prior, the $l_1$ prior, the total variation prior, and the deep image prior (DIP). We find that DIP outperforms the other three priors in terms of representational capability and visual performance. However, the performance of DIP deteriorates when the number of iterations exceeds a certain threshold due to overfitting. To address this issue, we propose a novel method (MCDIP-ADMM) based on Multi-Code Deep Image Prior and plug-and-play Alternative Direction Method of Multipliers. Specifically, MCDIP utilizes multiple latent codes to generate a series of feature maps at an intermediate layer within a generator model. These maps are then composed with trainable weights, representing the complete image prior. Experimental results demonstrate the superior performance of the proposed MCDIP-ADMM compared to three existing competitors. In the case of parallel beam projection with Gaussian noise, MCDIP-ADMM achieves an average improvement of 4.3 dB over DIP, 1.7 dB over ADMM DIP-WTV, and 1.2 dB over PnP-DIP in terms of PSNR. Similarly, for fan-beam projection with Poisson noise, MCDIP-ADMM achieves an average improvement of 3.09 dB over DIP, 1.86 dB over ADMM DIP-WTV, and 0.84 dB over PnP-DIP in terms of PSNR.

eess.IV↗