arXiv Science⌕ Search

arXiv subjects

Cheng Tan

Publications and source records attributed to Cheng Tan.

At least 163 records · Page 9Linked to original sources

Generative De Novo Protein Design with Global Context

The linear sequence of amino acids determines protein structure and function. Protein design, known as the inverse of protein structure prediction, aims to obtain a novel protein sequence that will fold into the defined structure. Recent works on computational protein design have studied designing sequences for the desired backbone structure with local positional information and achieved competitive performance. However, similar local environments in different backbone structures may result in different amino acids, indicating that protein structure's global context matters. Thus, we propose the Global-Context Aware generative de novo protein design method (GCA), consisting of local and global modules. While local modules focus on relationships between neighbor amino acids, global modules explicitly capture non-local contexts. Experimental results demonstrate that the proposed GCA method outperforms state-of-the-arts on de novo protein design. Our code and pretrained model will be released.

q-bio.BM↗

Imaging current control of magnetization in Fe$_3$GeTe$_2$ with a widefield nitrogen-vacancy microscope

Van der Waals (vdW) magnets are appealing candidates for realising spintronic devices that exploit current control of magnetization (e.g. switching or domain wall motion), but so far experimental demonstrations have been sparse, in part because of challenges associated with imaging the magnetization in these systems. Widefield nitrogen-vacancy (NV) microscopy allows rapid, quantitative magnetic imaging across entire vdW flakes, ideal for capturing changes in the micromagnetic structure due to an electric current. Here we use a widefield NV microscope to study the effect of current injection in thin flakes ($\sim10$nm) of the vdW ferromagnet Fe$_3$GeTe$_2$ (FGT). We first observe current-reduced coercivity on an individual domain level, where current injection in FGT causes substantial reduction in the magnetic field required to locally reverse the magnetisation. We then explore the possibility of current-induced domain-wall motion, and provide preliminary evidence for such a motion under relatively low current densities, suggesting the existence of strong current-induced torques in our devices. Our results illustrate the applicability of widefield NV microscopy to imaging spintronic phenomena in vdW magnets, highlight the possibility of efficient magnetization control by direct current injection without assistance from an adjacent conductor, and motivate further investigations of the effect of currents in FGT and other vdW magnets.

cond-mat.mes-hall↗

PrefixMol: Target- and Chemistry-aware Molecule Design via Prefix Embedding

Is there a unified model for generating molecules considering different conditions, such as binding pockets and chemical properties? Although target-aware generative models have made significant advances in drug design, they do not consider chemistry conditions and cannot guarantee the desired chemical properties. Unfortunately, merging the target-aware and chemical-aware models into a unified model to meet customized requirements may lead to the problem of negative transfer. Inspired by the success of multi-task learning in the NLP area, we use prefix embeddings to provide a novel generative model that considers both the targeted pocket's circumstances and a variety of chemical properties. All conditional information is represented as learnable features, which the generative model subsequently employs as a contextual prompt. Experiments show that our model exhibits good controllability in both single and multi-conditional molecular generation. The controllability enables us to outperform previous structure-based drug design methods. More interestingly, we open up the attention mechanism and reveal coupling relationships between conditions, providing guidance for multi-conditional molecule generation.

cs.AI↗

DiffSDS: A language diffusion model for protein backbone inpainting under geometric conditions and constraints

Have you ever been troubled by the complexity and computational cost of SE(3) protein structure modeling and been amazed by the simplicity and power of language modeling? Recent work has shown promise in simplifying protein structures as sequences of protein angles; therefore, language models could be used for unconstrained protein backbone generation. Unfortunately, such simplification is unsuitable for the constrained protein inpainting problem, where the model needs to recover masked structures conditioned on unmasked ones, as it dramatically increases the computing cost of geometric constraints. To overcome this dilemma, we suggest inserting a hidden \textbf{a}tomic \textbf{d}irection \textbf{s}pace (\textbf{ADS}) upon the language model, converting invariant backbone angles into equivalent direction vectors and preserving the simplicity, called Seq2Direct encoder ($\text{Enc}_{s2d}$). Geometric constraints could be efficiently imposed on the newly introduced direction space. A Direct2Seq decoder ($\text{Dec}_{d2s}$) with mathematical guarantees is also introduced to develop a \textbf{SDS} ($\text{Enc}_{s2d}$+$\text{Dec}_{d2s}$) model. We apply the SDS model as the denoising neural network during the conditional diffusion process, resulting in a constrained generative model--\textbf{DiffSDS}. Extensive experiments show that the plug-and-play ADS could transform the language model into a strong structural model without loss of simplicity. More importantly, the proposed DiffSDS outperforms previous strong baselines by a large margin on the task of protein inpainting.

q-bio.QM↗

NNSmith: Generating Diverse and Valid Test Cases for Deep Learning Compilers

Deep-learning (DL) compilers such as TVM and TensorRT are increasingly being used to optimize deep neural network (DNN) models to meet performance, resource utilization and other requirements. Bugs in these compilers can result in models whose semantics differ from the original ones, producing incorrect results that corrupt the correctness of downstream applications. However, finding bugs in these compilers is challenging due to their complexity. In this work, we propose a new fuzz testing approach for finding bugs in deep-learning compilers. Our core approach consists of (i) generating diverse yet valid DNN test models that can exercise a large part of the compiler's transformation logic using light-weight operator specifications; (ii) performing gradient-based search to find model inputs that avoid any floating-point exceptional values during model execution, reducing the chance of missed bugs or false alarms; and (iii) using differential testing to identify bugs. We implemented this approach in NNSmith which has found 72 new bugs for TVM, TensorRT, ONNXRuntime, and PyTorch to date. Of these 58 have been confirmed and 51 have been fixed by their respective project maintainers.

cs.LG↗

Protein Language Models and Structure Prediction: Connection and Progression

The prediction of protein structures from sequences is an important task for function prediction, drug design, and related biological processes understanding. Recent advances have proved the power of language models (LMs) in processing the protein sequence databases, which inherit the advantages of attention networks and capture useful information in learning representations for proteins. The past two years have witnessed remarkable success in tertiary protein structure prediction (PSP), including evolution-based and single-sequence-based PSP. It seems that instead of using energy-based models and sampling procedures, protein language model (pLM)-based pipelines have emerged as mainstream paradigms in PSP. Despite the fruitful progress, the PSP community needs a systematic and up-to-date survey to help bridge the gap between LMs in the natural language processing (NLP) and PSP domains and introduce their methodologies, advancements and practical applications. To this end, in this paper, we first introduce the similarities between protein and human languages that allow LMs extended to pLMs, and applied to protein databases. Then, we systematically review recent advances in LMs and pLMs from the perspectives of network architectures, pre-training strategies, applications, and commonly-used protein databases. Next, different types of methods for PSP are discussed, particularly how the pLM-based architectures function in the process of protein folding. Finally, we identify challenges faced by the PSP community and foresee promising research directions along with the advances of pLMs. This survey aims to be a hands-on guide for researchers to understand PSP methods, develop pLMs and tackle challenging problems in this field for practical purposes.

q-bio.QM↗

Leveraging Graph-based Cross-modal Information Fusion for Neural Sign Language Translation

Sign Language (SL), as the mother tongue of the deaf community, is a special visual language that most hearing people cannot understand. In recent years, neural Sign Language Translation (SLT), as a possible way for bridging communication gap between the deaf and the hearing people, has attracted widespread academic attention. We found that the current mainstream end-to-end neural SLT models, which tries to learning language knowledge in a weakly supervised manner, could not mine enough semantic information under the condition of low data resources. Therefore, we propose to introduce additional word-level semantic knowledge of sign language linguistics to assist in improving current end-to-end neural SLT models. Concretely, we propose a novel neural SLT model with multi-modal feature fusion based on the dynamic graph, in which the cross-modal information, i.e. text and video, is first assembled as a dynamic graph according to their correlation, and then the graph is processed by a multi-modal graph encoder to generate the multi-modal embeddings for further usage in the subsequent neural translation models. To the best of our knowledge, we are the first to introduce graph neural networks, for fusing multi-modal information, into neural sign language translation models. Moreover, we conducted experiments on a publicly available popular SLT dataset RWTH-PHOENIX-Weather-2014T. and the quantitative experiments show that our method can improve the model.

cs.CL↗

Target-aware Molecular Graph Generation

Generating molecules with desired biological activities has attracted growing attention in drug discovery. Previous molecular generation models are designed as chemocentric methods that hardly consider the drug-target interaction, limiting their practical applications. In this paper, we aim to generate molecular drugs in a target-aware manner that bridges biological activity and molecular design. To solve this problem, we compile a benchmark dataset from several publicly available datasets and build baselines in a unified framework. Building on the recent advantages of flow-based molecular generation models, we propose SiamFlow, which forces the flow to fit the distribution of target sequence embeddings in latent space. Specifically, we employ an alignment loss and a uniform loss to bring target sequence embeddings and drug graph embeddings into agreements while avoiding collapse. Furthermore, we formulate the alignment into a one-to-many problem by learning spaces of target sequence embeddings. Experiments quantitatively show that our proposed method learns meaningful representations in the latent space toward the target-aware molecular graph generation and provides an alternative approach to bridge biology and chemistry in drug discovery.

cs.LG↗

Intrinsic new properties of a quantum spin liquid

Quantum fluctuations are expected to lead to highly entangled spin-liquid states in certain two-dimensional spin-1/2 compounds. We have synthesized and measured thermodynamic properties and muon spin relaxation rates in the copper-based two-dimensional triangular-lattice spin liquids Lu$_3$Cu$_2$Sb$_3$O$_{14}$ and Lu$_3$CuZnSb$_3$O$_{14}$. The former is the least disordered of this kind discovered to date. Magnetic entropy generation at high temperatures has been ruled out after carefully correcting for the lattice specific heat. Surprisingly, roughly half of the magnetic entropy is missing down to temperatures of O(10$^{-3}$) the exchange energy, independent of magnetic field up to $gμ_B H \gtrsim k_BΘ_W$, where $Θ_W$ is the Weiss temperature. The magnetic specific heat divided by temperature $C_M(T)/T$ and muon spin relaxation rate $λ(T)$ are both temperature-independent at low temperatures, followed by logarithmic decreases with increasing temperature. This behavior can be simply characterized by scale-invariant time-dependent fluctuations with a single parameter. Since no cooperative effects due to impurities are observed, the measured properties are intrinsic. They are evidence that in Lu$_3$Cu$_2$Sb$_3$O$_{14}$ massive quantum fluctuations lead to either a gigantic specific heat peak from singlet excitations at very low temperatures or, perhaps less likely, an extensively degenerate possibly topological singlet ground state.

cond-mat.str-el↗

CoSP: Co-supervised pretraining of pocket and ligand

Can we inject the pocket-ligand interaction knowledge into the pre-trained model and jointly learn their chemical space? Pretraining molecules and proteins has attracted considerable attention in recent years, while most of these approaches focus on learning one of the chemical spaces and lack the injection of biological knowledge. We propose a co-supervised pretraining (CoSP) framework to simultaneously learn 3D pocket and ligand representations. We use a gated geometric message passing layer to model both 3D pockets and ligands, where each node's chemical features, geometric position and orientation are considered. To learn biological meaningful embeddings, we inject the pocket-ligand interaction knowledge into the pretraining model via contrastive loss. Considering the specificity of molecules, we further propose a chemical similarity-enhanced negative sampling strategy to improve the contrastive learning performance. Through extensive experiments, we conclude that CoSP can achieve competitive results in pocket matching, molecule property predictions, and virtual screening.

cs.LG↗

SimVP: Simpler yet Better Video Prediction

From CNN, RNN, to ViT, we have witnessed remarkable advancements in video prediction, incorporating auxiliary inputs, elaborate neural architectures, and sophisticated training strategies. We admire these progresses but are confused about the necessity: is there a simple method that can perform comparably well? This paper proposes SimVP, a simple video prediction model that is completely built upon CNN and trained by MSE loss in an end-to-end fashion. Without introducing any additional tricks and complicated strategies, we can achieve state-of-the-art performance on five benchmark datasets. Through extended experiments, we demonstrate that SimVP has strong generalization and extensibility on real-world datasets. The significant reduction of training cost makes it easier to scale to complex scenarios. We believe SimVP can serve as a solid baseline to stimulate the further development of video prediction. The code is available at \href{https://github.com/gaozhangyang/SimVP-Simpler-yet-Better-Video-Prediction}{Github}.

cs.CV↗

Hyperspherical Consistency Regularization

Recent advances in contrastive learning have enlightened diverse applications across various semi-supervised fields. Jointly training supervised learning and unsupervised learning with a shared feature encoder becomes a common scheme. Though it benefits from taking advantage of both feature-dependent information from self-supervised learning and label-dependent information from supervised learning, this scheme remains suffering from bias of the classifier. In this work, we systematically explore the relationship between self-supervised learning and supervised learning, and study how self-supervised learning helps robust data-efficient deep learning. We propose hyperspherical consistency regularization (HCR), a simple yet effective plug-and-play method, to regularize the classifier using feature-dependent information and thus avoid bias from labels. Specifically, HCR first projects logits from the classifier and feature projections from the projection head on the respective hypersphere, then it enforces data points on hyperspheres to have similar structures by minimizing binary cross entropy of pairwise distances' similarity metrics. Extensive experiments on semi-supervised and weakly-supervised learning demonstrate the effectiveness of our method, by showing superior performance with HCR.

cs.LG↗

Tunable and giant valley-selective Hall effect in gapped bilayer graphene

Berry curvature is analogous to magnetic field but in momentum space and is commonly present in materials with non-trivial quantum geometry. It endows Bloch electrons with transverse anomalous velocities to produce Hall-like currents even in the absence of a magnetic field. We report the direct observation of in situ tunable valley-selective Hall effect (VSHE), where inversion symmetry, and thus the geometric phase of electrons, is controllable by an out-of-plane electric field. We use high-quality bilayer graphene with an intrinsic and tunable bandgap, illuminated by circularly polarized mid-infrared light and confirm that the observed Hall voltage arises from an optically-induced valley population. Compared with molybdenum disulfide, we find orders of magnitude larger VSHE, attributed to the inverse scaling of the Berry curvature with bandgap. By monitoring the valley-selective Hall conductivity, we study Berry curvature's evolution with bandgap. This in situ manipulation of VSHE paves the way for topological and quantum geometric opto-electronic devices, such as more robust switches and detectors.

cond-mat.mes-hall↗

Dissipation-enabled hydrodynamic conductivity in a tunable bandgap semiconductor

Electronic transport in the regime where carrier-carrier collisions are the dominant scattering mechanism has taken on new relevance with the advent of ultraclean two-dimensional materials. Here we present a combined theoretical and experimental study of ambipolar hydrodynamic transport in bilayer graphene demonstrating that the conductivity is given by the sum of two Drude-like terms that describe relative motion between electrons and holes, and the collective motion of the electron-hole plasma. As predicted, the measured conductivity of gapless, charge-neutral bilayer graphene is sample- and temperature-independent over a wide range. Away from neutrality, the electron-hole conductivity collapses to a single curve, and a set of just four fitting parameters provides quantitative agreement between theory and experiment at all densities, temperatures, and gaps measured. This work validates recent theories for dissipation-enabled hydrodynamic conductivity and creates a link between semiconductor physics and the emerging field of viscous electronics.

cond-mat.mes-hall↗

Gate-tunable exchange bias effect in FePS3-Fe5GeTe2 van der Waals heterostructures

Electrical gate-manipulated exchange bias (EB) effect is a long-term goal for spintronics applications. Meanwhile, the emergence of van der Waals (vdW) magnetic heterostructures provides ideal platforms for the study of interlayer magnetic coupling. However, to date, the electrical gate-controlled EB effect has yet to be realized in vdW heterostructures. Here, for the first time, we realized electrically-controllable EB effects in a vdW antiferromagnetic (AFM)-ferromagnetic (FM) heterostructure, FePS3-Fe5GeTe2. For pristine FePS3-Fe5GeTe2 heterostructures, sizable EB effects can be generated due to the strong interface coupling, which also depend on the thickness of the ferromagnetic layers. By applying a solid protonic gate, the EB effects can be electrically tuned largely by proton intercalations and deintercalations. The EB field reaches up to 23% of the coercive field and the blocking temperature exceeds 50 K at Vg= -3.15 V. The proton intercalations not only tune the average magnetic exchange coupling, but also change the AFM configurations and transform the heterointerface between an uncompensated AFM-FM interface and a compensated AFM-FM interface. These alterations result in a dramatic modulation of the total interface exchange coupling and the resultant EB effects. The study is a significant step towards vdW heterostructure-based magnetic logic for future low-energy electronics.

cond-mat.mtrl-sci↗

Electrically controlled superconductor-insulator transition and giant anomalous Hall effect in kagome metal CsV3Sb5 nanoflakes

The electronic correlations (e.g. unconventional superconductivity (SC), chiral charge order and nematic order) and giant anomalous Hall effect (AHE) in topological kagome metals AV3Sb5 (A= K, Rb, and Cs) have attracted great interest. Electrical control of those correlated electronic states and AHE allows us to resolve their own nature and origin and to discover new quantum phenomena. Here, we show that a protonic gate can largely modulate the effective disorders and carrier density in CsV3Sb5 nanoflakes, leading to significant modifications of SC, unusual charge density wave (CDW) and giant AHE. Notably, we observed a direct superconductor-insulator transition (SIT) driven by superconducting phase fluctuation due to the doping-enhanced disorders, in addition to a large suppression of CDW. Meanwhile, the carrier density modulation shifts the Fermi level across the CDW gap and gives rise to a nontrivial evolution of AHE, in line with the asymmetric density of states of CDW sub-bands near the saddle point. With the first-principles calculations, we suggest the extrinsic skew scattering of holes in the nearly flat bands with finite Berry curvature by multiple impurities accounts for the giant AHE. Our work uncovers a disorder-driven bosonic SIT, outlines a global picture of the giant AHE and reveals its correlation with the unconventional CDW in the AV3Sb5 family.

cond-mat.mtrl-sci↗

I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through Islandization

Graph Convolutional Networks (GCNs) have drawn tremendous attention in the past three years. Compared with other deep learning modalities, high-performance hardware acceleration of GCNs is as critical but even more challenging. The hurdles arise from the poor data locality and redundant computation due to the large size, high sparsity, and irregular non-zero distribution of real-world graphs. In this paper we propose a novel hardware accelerator for GCN inference, called I-GCN, that significantly improves data locality and reduces unnecessary computation. The mechanism is a new online graph restructuring algorithm we refer to as islandization. The proposed algorithm finds clusters of nodes with strong internal but weak external connections. The islandization process yields two major benefits. First, by processing islands rather than individual nodes, there is better on-chip data reuse and fewer off-chip memory accesses. Second, there is less redundant computation as aggregation for common/shared neighbors in an island can be reused. The parallel search, identification, and leverage of graph islands are all handled purely in hardware at runtime working in an incremental pipeline. This is done without any preprocessing of the graph data or adjustment of the GCN model structure. Experimental results show that I-GCN can significantly reduce off-chip accesses and prune 38% of aggregation operations, leading to performance speedups over CPUs, GPUs, the prior art GCN accelerators of 5549x, 403x, and 5.7x on average, respectively.

cs.AR↗

AlphaDesign: A graph protein design method and benchmark on AlphaFoldDB

While DeepMind has tentatively solved protein folding, its inverse problem -- protein design which predicts protein sequences from their 3D structures -- still faces significant challenges. Particularly, the lack of large-scale standardized benchmark and poor accuray hinder the research progress. In order to standardize comparisons and draw more research interest, we use AlphaFold DB, one of the world's largest protein structure databases, to establish a new graph-based benchmark -- AlphaDesign. Based on AlphaDesign, we propose a new method called ADesign to improve accuracy by introducing protein angles as new features, using a simplified graph transformer encoder (SGT), and proposing a confidence-aware protein decoder (CPD). Meanwhile, SGT and CPD also improve model efficiency by simplifying the training and testing procedures. Experiments show that ADesign significantly outperforms previous graph models, e.g., the average accuracy is improved by 8\%, and the inference speed is 40+ times faster than before.

q-bio.QM↗