arXiv Science⌕ Search

arXiv subjects

Cheng Tan

Publications and source records attributed to Cheng Tan.

At least 109 records · Page 6Linked to original sources

PipeLLM: Fast and Confidential Large Language Model Services with Speculative Pipelined Encryption

Confidential computing on GPUs, like NVIDIA H100, mitigates the security risks of outsourced Large Language Models (LLMs) by implementing strong isolation and data encryption. Nonetheless, this encryption incurs a significant performance overhead, reaching up to 52.8 percent and 88.2 percent throughput drop when serving OPT-30B and OPT-66B, respectively. To address this challenge, we introduce PipeLLM, a user-transparent runtime system. PipeLLM removes the overhead by overlapping the encryption and GPU computation through pipelining - an idea inspired by the CPU instruction pipelining - thereby effectively concealing the latency increase caused by encryption. The primary technical challenge is that, unlike CPUs, the encryption module lacks prior knowledge of the specific data needing encryption until it is requested by the GPUs. To this end, we propose speculative pipelined encryption to predict the data requiring encryption by analyzing the serving patterns of LLMs. Further, we have developed an efficient, low-cost pipeline relinquishing approach for instances of incorrect predictions. Our experiments on NVIDIA H100 GPU show that compared with vanilla systems without confidential computing (e.g., vLLM, PEFT, and FlexGen), PipeLLM incurs modest overhead (less than 19.6 percent in throughput) across various LLM sizes, from 13B to 175B.

cs.CR↗

GATES: Graph Attention Network with Global Expression Fusion for Deciphering Spatial Transcriptome Architectures

Single-cell spatial transcriptomics (ST) offers a unique approach to measuring gene expression profiles and spatial cell locations simultaneously. However, most existing ST methods assume that cells in closer spatial proximity exhibit more similar gene expression patterns. Such assumption typically results in graph structures that prioritize local spatial information while overlooking global patterns, limiting the ability to fully capture the broader structural features of biological tissues. To overcome this limitation, we propose GATES (Graph Attention neTwork with global Expression fuSion), a novel model designed to capture structural details in spatial transcriptomics data. GATES first constructs an expression graph that integrates local and global information by leveraging both spatial proximity and gene expression similarity. The model then employs an autoencoder with adaptive attention to assign proper weights for neighboring nodes, enhancing its capability of feature extraction. By fusing features of both the spatial and expression graphs, GATES effectively balances spatial context with gene expression data. Experimental results across multiple datasets demonstrate that GATES significantly outperforms existing methods in identifying spatial domains, highlighting its potential for analyzing complex biological tissues. Our code can be accessed on GitHub at https://github.com/xiaoxiongtao/GATES.

q-bio.GN↗

Imaging magnetic switching in orthogonally twisted stacks of a van der Waals antiferromagnet

Stacking van der Waals magnets holds promise for creating new hybrid materials with properties that do not exist in bulk materials. Here we investigate orthogonally twisted stacks of the van der Waals antiferromagnet CrSBr, aiming to exploit an extreme misalignment of magnetic anisotropy across the twisted interface.Using nitrogen-vacancy centre microscopy, we construct vector maps of the magnetisation, and track their evolution under an external field, in a range of twisted compensated and uncompensated configurations differing by the number of layers. We show that twisted stacking consistently modifies the local magnetic switching behaviour of constituent flakes, and that these modifications are spatially non-uniform. In the case of compensated component flakes (even number of layers), we demonstrate that the combination of dipolar coupling and stacking-induced strain can reduce the switching field by over an order of magnitude. Conversely, in uncompensated component flakes (odd number of layers), we observe indications of a non-zero interlayer exchange interaction between twisted flakes during magnetization reversal, which can persistently modify magnetic order. This work highlights the importance of spatial imaging in investigating stacking-induced magnetic effects, particularly in the case of twistronics where spatial variation is expected and can be conflated with structural imperfections.

cond-mat.mes-hall↗

FlexMol: A Flexible Toolkit for Benchmarking Molecular Relational Learning

Molecular relational learning (MRL) is crucial for understanding the interaction behaviors between molecular pairs, a critical aspect of drug discovery and development. However, the large feasible model space of MRL poses significant challenges to benchmarking, and existing MRL frameworks face limitations in flexibility and scope. To address these challenges, avoid repetitive coding efforts, and ensure fair comparison of models, we introduce FlexMol, a comprehensive toolkit designed to facilitate the construction and evaluation of diverse model architectures across various datasets and performance metrics. FlexMol offers a robust suite of preset model components, including 16 drug encoders, 13 protein sequence encoders, 9 protein structure encoders, and 7 interaction layers. With its easy-to-use API and flexibility, FlexMol supports the dynamic construction of over 70, 000 distinct combinations of model architectures. Additionally, we provide detailed benchmark results and code examples to demonstrate FlexMol's effectiveness in simplifying and standardizing MRL model development and comparison.

cs.LG↗

CBGBench: Fill in the Blank of Protein-Molecule Complex Binding Graph

Structure-based drug design (SBDD) aims to generate potential drugs that can bind to a target protein and is greatly expedited by the aid of AI techniques in generative models. However, a lack of systematic understanding persists due to the diverse settings, complex implementation, difficult reproducibility, and task singularity. Firstly, the absence of standardization can lead to unfair comparisons and inconclusive insights. To address this dilemma, we propose CBGBench, a comprehensive benchmark for SBDD, that unifies the task as a generative heterogeneous graph completion, analogous to fill-in-the-blank of the 3D complex binding graph. By categorizing existing methods based on their attributes, CBGBench facilitates a modular and extensible framework that implements various cutting-edge methods. Secondly, a single task on \textit{de novo} molecule generation can hardly reflect their capabilities. To broaden the scope, we have adapted these models to a range of tasks essential in drug design, which are considered sub-tasks within the graph fill-in-the-blank tasks. These tasks include the generative designation of \textit{de novo} molecules, linkers, fragments, scaffolds, and sidechains, all conditioned on the structures of protein pockets. Our evaluations are conducted with fairness, encompassing comprehensive perspectives on interaction, chemical properties, geometry authenticity, and substructure validity. We further provide the pre-trained versions of the state-of-the-art models and deep insights with analysis from empirical studies. The codebase for CBGBench is publicly accessible at \url{https://github.com/Edapinenut/CBGBench}.

cs.LG↗

Muon Spin Relaxation Study of Spin Dynamics on a Kitaev honeycomb material H$_3$LiIr$_2$O$_6$

The vacancy effect in quantum spin liquid (QSL) has been extensively studied. A finite density of random vacancies in the Kitaev model can lead to a pileup of low-energy density of states (DOS), which is generally experimentally determined by a scaling behavior of thermodynamic or magnetization quantities. Here, we report detailed muon spin relaxation ($μ$SR) results of H$_3$LiIr$_2$O$_6$, a Kitaev QSL candidate with vacancies. The absence of magnetic order is confirmed down to 80 mK, and the spin fluctuations are found to be persistent at low temperatures. Intriguingly, the time-field scaling law of longitudinal-field (LF)-$μ$SR polarization is observed down to 0.1 K. This indicates a dynamical scaling, whose critical exponent 0.46 is excellently consistent with the scaling behavior of specific heat and magnetization data. All the observations point to the finite DOS with the form $N(E) \sim E^{-0.5}$ , which is expected for the Kitaev QSL in the presence of vacancies. Our μSR study provides a dynamical fingerprint of the power-law low-energy DOS, and introduces a crucial new insight into the vacancy effect in QSL.

cond-mat.str-el↗

High-speed ultra-broadband detection based on interfacial work function internal photoemission detector

High-speed ultra-broadband detectors play a crucial role in aerospace technology, and national security etc. The interfacial work function internal photoemission (IWIP) detector employs multiple absorption mechanism comprehensively across different wavelength band to achieve complete photon type detection, which makes it possible to realize high-speed and ultra-broadband simultaneously. We propose a ratchet heterojunction IWIP (HEIWIP) detector, which shows 3-165THz ultra-broadband coverage. The high-speed response is investigated in detail by both microwave rectification technology and high-speed modulated terahertz light. Up to 5.1GHz 3dB bandwidth is acquired in terms of microwave rectification measurement. And 4.255GHz inter-mode optical beat note signal was successfully detected.

physics.ins-det↗

OpenMixup: Open Mixup Toolbox and Benchmark for Visual Representation Learning

Mixup augmentation has emerged as a widely used technique for improving the generalization ability of deep neural networks (DNNs). However, the lack of standardized implementations and benchmarks has impeded recent progress, resulting in poor reproducibility, unfair comparisons, and conflicting insights. In this paper, we introduce OpenMixup, the first mixup augmentation codebase, and benchmark for visual representation learning. Specifically, we train 18 representative mixup baselines from scratch and rigorously evaluate them across 11 image datasets of varying scales and granularity, ranging from fine-grained scenarios to complex non-iconic scenes. We also open-source our modular codebase, including a collection of popular vision backbones, optimization strategies, and analysis toolkits, which not only supports the benchmarking but enables broader mixup applications beyond classification, such as self-supervised learning and regression tasks. Through experiments and empirical analysis, we gain observations and insights on mixup performance-efficiency trade-offs, generalization, and optimization behaviors, and thereby identify preferred choices for different needs. To the best of our knowledge, OpenMixup has facilitated several recent studies. We believe this work can further advance reproducible mixup augmentation research and thereby lay a solid ground for future progress in the community. The source code and user documents are available at \url{https://github.com/Westlake-AI/openmixup}.

cs.CV↗

Switch EMA: A Free Lunch for Better Flatness and Sharpness

Exponential Moving Average (EMA) is a widely used weight averaging (WA) regularization to learn flat optima for better generalizations without extra cost in deep neural network (DNN) optimization. Despite achieving better flatness, existing WA methods might fall into worse final performances or require extra test-time computations. This work unveils the full potential of EMA with a single line of modification, i.e., switching the EMA parameters to the original model after each epoch, dubbed as Switch EMA (SEMA). From both theoretical and empirical aspects, we demonstrate that SEMA can help DNNs to reach generalization optima that better trade-off between flatness and sharpness. To verify the effectiveness of SEMA, we conduct comparison experiments with discriminative, generative, and regression tasks on vision and language datasets, including image classification, self-supervised learning, object detection and segmentation, image generation, video prediction, attribute regression, and language modeling. Comprehensive results with popular optimizers and networks show that SEMA is a free lunch for DNN training by improving performances and boosting convergence speeds.

cs.LG↗

Learning to Model Graph Structural Information on MLPs via Graph Structure Self-Contrasting

Recent years have witnessed great success in handling graph-related tasks with Graph Neural Networks (GNNs). However, most existing GNNs are based on message passing to perform feature aggregation and transformation, where the structural information is explicitly involved in the forward propagation by coupling with node features through graph convolution at each layer. As a result, subtle feature noise or structure perturbation may cause severe error propagation, resulting in extremely poor robustness. In this paper, we rethink the roles played by graph structural information in graph data training and identify that message passing is not the only path to modeling structural information. Inspired by this, we propose a simple but effective Graph Structure Self-Contrasting (GSSC) framework that learns graph structural information without message passing. The proposed framework is based purely on Multi-Layer Perceptrons (MLPs), where the structural information is only implicitly incorporated as prior knowledge to guide the computation of supervision signals, substituting the explicit message propagation as in GNNs. Specifically, it first applies structural sparsification to remove potentially uninformative or noisy edges in the neighborhood, and then performs structural self-contrasting in the sparsified neighborhood to learn robust node representations. Finally, structural sparsification and self-contrasting are formulated as a bi-level optimization problem and solved in a unified framework. Extensive experiments have qualitatively and quantitatively demonstrated that the GSSC framework can produce truly encouraging performance with better generalization and robustness than other leading competitors.

cs.LG↗

Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-Consistency Training

Multimodal reasoning is a challenging task that requires models to reason across multiple modalities to answer questions. Existing approaches have made progress by incorporating language and visual modalities into a two-stage reasoning framework, separating rationale generation from answer inference. However, these approaches often fall short due to the inadequate quality of the generated rationales. In this work, we delve into the importance of rationales in model reasoning. We observe that when rationales are completely accurate, the model's accuracy significantly improves, highlighting the need for high-quality rationale generation. Motivated by this, we propose MC-CoT, a self-consistency training strategy that generates multiple rationales and answers, subsequently selecting the most accurate through a voting process. This approach not only enhances the quality of generated rationales but also leads to more accurate and robust answers. Through extensive experiments, we demonstrate that our approach significantly improves model performance across various benchmarks. Remarkably, we show that even smaller base models, when equipped with our proposed approach, can achieve results comparable to those of larger models, illustrating the potential of our approach in harnessing the power of rationales for improved multimodal reasoning. The code is available at https://github.com/chengtan9907/mc-cot.

cs.AI↗

Multi-species optically addressable spin defects in a van der Waals material

Optically addressable spin defects hosted in two-dimensional van der Waals materials represent a new frontier for quantum technologies, promising to lead to a new class of ultrathin quantum sensors and simulators. Recently, hexagonal boron nitride (hBN) has been shown to host several types of optically addressable spin defects, thus offering a unique opportunity to simultaneously address and utilise various spin species in a single material. Here we demonstrate an interplay between two separate spin species within a single hBN crystal, namely $S=1$ boron vacancy defects and visible emitter spins. We unambiguously prove that the visible emitters are $S=\frac{1}{2}$ spins and further demonstrate room temperature coherent control and optical readout of both spin species. Importantly, by tuning the two spin species into resonance with each other, we observe cross-relaxation indicating strong inter-species dipolar coupling. We then demonstrate magnetic imaging using the $S=\frac{1}{2}$ defects, both under ambient and cryogenic conditions, and leverage their lack of intrinsic quantization axis to determine the anisotropic magnetic susceptibility of a test sample. Our results establish hBN as a versatile platform for quantum technologies in a van der Waals host at room temperature.

cond-mat.mes-hall↗

Deciphering RNA Secondary Structure Prediction: A Probabilistic K-Rook Matching Perspective

The secondary structure of ribonucleic acid (RNA) is more stable and accessible in the cell than its tertiary structure, making it essential for functional prediction. Although deep learning has shown promising results in this field, current methods suffer from poor generalization and high complexity. In this work, we reformulate the RNA secondary structure prediction as a K-Rook problem, thereby simplifying the prediction process into probabilistic matching within a finite solution space. Building on this innovative perspective, we introduce RFold, a simple yet effective method that learns to predict the most matching K-Rook solution from the given sequence. RFold employs a bi-dimensional optimization strategy that decomposes the probabilistic matching problem into row-wise and column-wise components to reduce the matching complexity, simplifying the solving process while guaranteeing the validity of the output. Extensive experiments demonstrate that RFold achieves competitive performance and about eight times faster inference efficiency than the state-of-the-art approaches. The code and Colab demo are available in (http://github.com/A4Bio/RFold).

q-bio.BM↗

Massive Dirac Fermions and Strong Shubnikov-de Haas Oscillations in Topological Insulator Sm,Fe:Bi2Se3 Single Crystals

Topological insulators (TIs) are emergent materials with unique band structure, which allow the study of quantum effect in solids, as well as contribute to high performance quantum devices. To achieve the better performance of TI, here we present a co-doping strategy using synergistic rare-earth Sm and transition-metal Fe dopants in Bi2Se3 single crystals, which combine the advantages of both transition metal doped TI (high ferromagnetic ordering temperature and observed QAHE), and rare-earth doped TI (large magnetic moments and significant spin orbit coupling). In the as-grown single crystals, clear evidences of ferromagnetic ordering were observed. The angle resolve photoemission spectroscopy indicate the ferromagnetism opens a 44 meV band gap at surface Dirac point. Moreover, the carrier mobility at 3 K is ~ 7400 cm2/Vs, and we thus observed an ultra-strong Shubnikov-de Haas oscillation in the longitudinal resistivity, as well as the Hall steps in transverse resistivity below 14 T. Our transport and angular resolved photoemission spectroscopy results suggest that the rare-earth and transition metal co-doping in Bi2Se3 system is a promising avenue implement the quantum anomalous Hall effect, as well as harnessing the massive Dirac fermion in electrical devices.

cond-mat.mtrl-sci↗

FoldToken2: Learning compact, invariant and generative protein structure language

The equivalent nature of 3D coordinates has posed long term challenges in protein structure representation learning, alignment, and generation. Can we create a compact and invariant language that equivalently represents protein structures? Towards this goal, we propose FoldToken2 to transfer equivariant structures into discrete tokens, while maintaining the recoverability of the original structures. From FoldToken1 to FoldToken2, we improve three key components: (1) invariant structure encoder, (2) vector-quantized compressor, and (3) equivalent structure decoder. We evaluate FoldToken2 on the protein structure reconstruction task and show that it outperforms previous FoldToken1 by 20\% in TMScore and 81\% in RMSD. FoldToken2 probably be the first method that works well on both single-chain and multi-chain protein structures quantization. We believe that FoldToken2 will inspire further improvement in protein structure representation learning, structure alignment, and structure generation tasks.

q-bio.BM↗

Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions

Large Language Models (LLMs) have demonstrated wide-ranging applications across various fields and have shown significant potential in the academic peer-review process. However, existing applications are primarily limited to static review generation based on submitted papers, which fail to capture the dynamic and iterative nature of real-world peer reviews. In this paper, we reformulate the peer-review process as a multi-turn, long-context dialogue, incorporating distinct roles for authors, reviewers, and decision makers. We construct a comprehensive dataset containing over 26,841 papers with 92,017 reviews collected from multiple sources, including the top-tier conference and prestigious journal. This dataset is meticulously designed to facilitate the applications of LLMs for multi-turn dialogues, effectively simulating the complete peer-review process. Furthermore, we propose a series of metrics to evaluate the performance of LLMs for each role under this reformulated peer-review setting, ensuring fair and comprehensive evaluations. We believe this work provides a promising perspective on enhancing the LLM-driven peer-review process by incorporating dynamic, role-based interactions. It aligns closely with the iterative and interactive nature of real-world academic peer review, offering a robust foundation for future research and development in this area. We open-source the dataset at https://github.com/chengtan9907/ReviewMT.

cs.CL↗

GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models

The Genomic Foundation Model (GFM) paradigm is expected to facilitate the extraction of generalizable representations from massive genomic data, thereby enabling their application across a spectrum of downstream applications. Despite advancements, a lack of evaluation framework makes it difficult to ensure equitable assessment due to experimental settings, model intricacy, benchmark datasets, and reproducibility challenges. In the absence of standardization, comparative analyses risk becoming biased and unreliable. To surmount this impasse, we introduce GenBench, a comprehensive benchmarking suite specifically tailored for evaluating the efficacy of Genomic Foundation Models. GenBench offers a modular and expandable framework that encapsulates a variety of state-of-the-art methodologies. Through systematic evaluations of datasets spanning diverse biological domains with a particular emphasis on both short-range and long-range genomic tasks, firstly including the three most important DNA tasks covering Coding Region, Non-Coding Region, Genome Structure, etc. Moreover, We provide a nuanced analysis of the interplay between model architecture and dataset characteristics on task-specific performance. Our findings reveal an interesting observation: independent of the number of parameters, the discernible difference in preference between the attention-based and convolution-based models on short- and long-range tasks may provide insights into the future design of GFM.

q-bio.GN↗

VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence Modeling

Similar to natural language models, pre-trained genome language models are proposed to capture the underlying intricacies within genomes with unsupervised sequence modeling. They have become essential tools for researchers and practitioners in biology. However, the hand-crafted tokenization policies used in these models may not encode the most discriminative patterns from the limited vocabulary of genomic data. In this paper, we introduce VQDNA, a general-purpose framework that renovates genome tokenization from the perspective of genome vocabulary learning. By leveraging vector-quantized codebooks as learnable vocabulary, VQDNA can adaptively tokenize genomes into pattern-aware embeddings in an end-to-end manner. To further push its limits, we propose Hierarchical Residual Quantization (HRQ), where varying scales of codebooks are designed in a hierarchy to enrich the genome vocabulary in a coarse-to-fine manner. Extensive experiments on 32 genome datasets demonstrate VQDNA's superiority and favorable parameter efficiency compared to existing genome language models. Notably, empirical analysis of SARS-CoV-2 mutations reveals the fine-grained pattern awareness and biological significance of learned HRQ vocabulary, highlighting its untapped potential for broader applications in genomics.

q-bio.GN↗