arXiv ScienceSearch

arXiv subjects

Mu Qiao

Publications and source records attributed to Mu Qiao.

At least 19 recordsLinked to original sources

TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents

GUI agents accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usage. Training-free visual token pruning can reduce this cost, but cache reuse introduces a fundamental constraint. Once tokens are discarded, the corresponding visual evidence cannot be recovered without re-encoding. Pruning therefore becomes an \textit{irreversible admission decision} that must remain useful for unknown future targets while preserving coverage of operable regions under tight budgets. To address these challenges, we propose \textbf{\method{}}, a training-free framework for \emph{\textbf{T}rajectory-\textbf{r}obust \textbf{A}dmission and \textbf{C}overage-aware \textbf{E}vidence ordering}. Specifically, we combine a query-independent layout-derived interaction prior with instruction relevance and feature novelty to rank visual evidence according to both potential future utility and diversity. Then, we reserve part of the budget for native visual tokens distributed across the screen, repairing missing spatial coverage without breaking the ordering. Together, these mechanisms produce a nested token order, allowing retained visual evidence to shrink monotonically across budgets while remaining reusable throughout the trajectory. Finally, our monotone KV contraction incrementally contracts retired frames into compact session state, avoiding repeated visual encoding or pruning. Extensive experiments across six GUI benchmarks and diverse models verify the effectiveness of our proposed \method{} under tight budgets. The source code will be released.

cs.CV

Deterministic atom-shuttle interconnects via ultrafast atom-ion entangling gate

Neutral-atom arrays and trapped-ion crystals offer complementary strengths for fault-tolerant quantum computing but lack a fast way to deterministically interact. Here we propose a controlled-$Z$ gate generated by the charge-induced-dipole ($C_4$) force between a Rydberg-excited atom and a trapped ion, balanced by a spin-dependent optical Magnus force on the ion that closes phase-space trajectories within a few microseconds. Toggling the Rydberg state extends the scheme to multi-ion crystals at negligible overhead. The resulting ${\sim}5\,$kHz atom shuttle accelerates short-distance QCCD links and enables hybrid qLDPC memories in which atom logical qubits are written onto an ion block treated as a passive storage zone. We perform circuit-level Monte Carlo simulations and find that the hybrid architecture supports orders of magnitude more operations than atom-only or ion-only architectures at fixed code distance and logical error rate.

quant-ph

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs

Multimodal Large Language Models (MLLMs) incur prohibitive inference costs due to long visual token sequences. Training-free visual token reduction provides an efficient solution. However, existing methods distort attention distributions, giving rise to a phenomenon we term Attention Logit Collapse. To address this issue, we propose ERA, an Entropy-guided visual token pruning framework with Rectified Attention for efficient MLLMs. Specifically, ERA comprises three crucial components: Dual-view Entropy Pruning (DEP), Bias-aware Token Recycling (BTR), and Logit-preserving Attention Rectification (LAR). First, DEP identifies representative anchor tokens by jointly modeling visual diversity and head-wise saliency. BTR then recycles pruned tokens into their corresponding anchors while estimating a cluster-level logit bias. Building upon this, LAR injects the estimated bias into attention logits, effectively rectifying the collapse induced by token reduction. Together, these components preserve visual evidence even under aggressive compression, enabling robust performance across single-image, multi-image, and video settings on a wide range of MLLMs. Beyond delivering practical acceleration, ERA establishes logit-preserving visual token pruning as a principled framework for efficient MLLMs, unifying theoretical foundation, algorithmic design, and practical deployment. The code is at https://github.com/924973292/ERA.

cs.CV

Measuring spectral functions of doped magnets with Rydberg tweezer arrays

Spectroscopic measurements of single-particle spectral functions provide crucial insight into strongly correlated quantum matter by resolving the energy and spatial structure of elementary excitations. Here we introduce a spectroscopic protocol for single-charge injection with simultaneous spatial and energy resolution in a Rydberg tweezer array, effectively emulating scanning tunneling microscopy. By combining this protocol with single-atom-resolved imaging, we go beyond conventional spectroscopy by not only measuring the single-particle spectral function but also directly imaging the microscopic structure of the excitations underlying spectral resonances in frustrated $tJ$ Hamiltonians. We reveal resonances associated with the formation of bound magnetic polarons -- composite quasiparticles consisting of a mobile hole bound to a magnon -- and directly extract their binding energy, spatial extent, and spin character. Finally, by exploiting the spatial tunability of our platform, we measure the local density of states across different lattice geometries. Our work establishes Rydberg tweezer arrays as a powerful platform for spectroscopic studies of strongly correlated models, offering microscopic control and direct real-space access to emergent quasiparticles in engineered quantum matter.

cond-mat.quant-gas

Dirac Spin Liquid Candidate in a Rydberg Quantum Simulator

We experimentally investigate a frustrated spin-exchange antiferromagnet in a quantum simulator, composed of N = 114 dipolar Rydberg atoms arranged into a kagome array. Motivated by a recent theoretical proposal of a gapless U(1) Dirac spin liquid ground state, we use local addressing to adiabatically prepare low-energy states. We measure the local polarization and spin-spin correlations over this adiabatic protocol, and observe our system move from a staggered product state, through an intermediate magnetic crystal, and finally into a disordered, correlated liquid. We estimate the entropy density of this atomic liquid to be similar to that of frustrated magnetic insulators at liquid nitrogen temperatures. We compare the correlations in our liquid to those of a simple, parameter-free ansatz for the Dirac spin liquid, and find good agreement in the sign structure and spatial decay. Finally, we probe the static susceptibility of our system to a local field perturbation and to a geometrical distortion. Our results establish Rydberg atom arrays as a promising platform for the preparation and microscopic characterization of quantum spin liquid candidates.

cond-mat.quant-gas

Kinetically-induced bound states in a frustrated Rydberg tweezer array

Understanding how particles bind into composite objects is a ubiquitous theme in physics, from the formation of molecules to hadrons in quantum chromodynamics and the pairing of charge carriers in superconductors. The formation of bound states usually originates from attractive interactions between particles. However, the binding can also arise purely from the motion of dopants due to kinetic frustration, which is potentially related to unconventional pairing in moir\'e materials. Here, we report the first direct observation of kinetically-induced bound states between holes and magnons using a Rydberg atom array quantum simulator of the bosonic $t$-$J$ model in frustrated ladders and 2D lattices. First, we demonstrate the formation of mobile one-hole-one-magnon bound states. We then construct three-particle one-hole-two-magnon bound states and reveal the underlying binding mechanism by observing kinetically-induced singlet correlations. Finally, we investigate how mobile dopants structure their magnetic environment in a spin-balanced 2D triangular lattice, showing that a hole induces $120^\circ$ antiferromagnetic order, while a doublon dopant generates in-plane ferromagnetic correlations. Our results demonstrates compelling evidence of kinetically-induced binding, opening a new avenue to understand novel pairing mechanisms in correlated quantum materials like superconductors in moir\'e superlattices.

quant-ph

Probing spin-motion coupling of two Rydberg atoms by a Stern-Gerlach-like experiment

We propose and implement a protocol to measure the state-dependent motion of Rydberg atoms induced by dipole-dipole interactions. Our setup enables simultaneous readout of both the atomic internal state and position on a one-dimensional array of optical tweezers. We benchmark the protocol using two atoms in the same Rydberg state, which experience van der Waals repulsion, and measure velocities in agreement with theoretical predictions. When preparing the atoms in a different pair state, we observe an oscillatory dynamics that we attribute to the proximity of a macrodimer bound state. Finally, we perform a Stern-Gerlach-like experiment in which a superposition of the two previous pair states results in the separation of the atomic wavepacket into two macroscopically distinct trajectories, thereby demonstrating spin-motion coupling mediated by the interactions.

physics.atom-ph

Attention as Conditioning: What Classical Learning Theory Predicts About Linear Transformers

Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave. Predictive theories do exist, but in the literature on animal learning. We show that the state updates of the major linear-attention families are term-for-term identical with named models from a century of animal learning theory: linear attention implements Hebbian contiguity, DeltaNet implements Rescorla--Wagner error correction, and decay variants such as RetNet implement contiguity with a stimulus trace. This dictionary turns conditioning phenomena into testable statements about the in-context behavior of linear transformers, while distinguishing algebraic consequences from empirical measurements. Algebraically, it yields an exact closed form for Kamin blocking, verified in simulation to $<10^{-7}$ across five learning rates. Empirically, it predicts a dissociation that survives training on generic in-context association: error-correcting attention exhibits cue competition, whereas contiguity-based attention does not. A single state also has two capacity regimes, with measured scaling exponents of 1.22 for faithful retrieval and 1.89 for identification, consistent with linear and near-quadratic predictions. Across the full head grid, retrieval error is governed primarily by total state size rather than its partition across heads, indicating that heads provide capacity rather than redundant copies. We also prove no spontaneous recovery for the analyzed single-state recurrences under cue-orthogonal retention trials; with a never-presented-cue control and probes within the trained positional range, we likewise find no recovery in trained models. Finally, we introduce PH-attention, a Pearce--Hall-inspired rule with an explicit feature-indexed associability state that yields cue-dependent learning rates and is absent from the token-computed gates we compare.

cs.LG

Unsupervised Evolutionary Cell Type Matching via Entropy-Minimized Optimal Transport

Identifying evolutionary correspondences between cell types across species is a fundamental challenge in comparative genomics and evolutionary biology. Existing approaches often rely on either reference-based matching, which imposes asymmetry by designating one species as the reference, or projection-based matching, which may increase computational complexity and obscure biological interpretability at the cell-type level. Here, we present OT-MESH, an unsupervised computational framework leveraging entropy-regularized optimal transport (OT) to systematically determine cross-species cell type homologies. Our method uniquely integrates the Minimize Entropy of Sinkhorn (MESH) technique to refine the OT plan, transforming diffuse transport matrices into sparse, interpretable correspondences. Through systematic evaluation on synthetic datasets, we demonstrate that OT-MESH achieves near-optimal matching accuracy with computational efficiency, while maintaining remarkable robustness to noise. Compared to other OT-based methods like RefCM, OT-MESH provides speedup while achieving comparable accuracy. Applied to retinal bipolar cells (BCs) and retinal ganglion cells (RGCs) from mouse and macaque, OT-MESH accurately recovers known evolutionary relationships and uncovers novel correspondences, one of which was independently validated experimentally. Thus, our framework offers a principled, scalable, and interpretable solution for evolutionary cell type mapping, facilitating deeper insights into cellular specialization and conservation across species.

q-bio.QM

Benchmarking direct and indirect dipolar spin-exchange interactions between two Rydberg atoms

We report on the experimental characterization of various types of spin-exchange interactions between two individual atoms, where pseudo-spin degrees of freedom are encoded in different Rydberg states. For the case of the direct dipole-dipole interaction between states of opposite parity, such as between $nS$ and $nP$, we investigate the effects of positional disorder arising from the residual atomic motion, on the coherence of spin-exchange oscillations. We then characterize an indirect dipolar spin exchange, i.e., the off-diagonal part of the van der Waals effective Hamiltonian that couples the states $nS$ and $(n+1)S$. Finally, we report on the observation of a new type of dipolar coupling, made resonant using addressable light-shifts and involving four different Rydberg levels: this exchange process is akin to electrically induced F\"orster resonance, but featuring local control. It exhibits an angular dependence distinct from the usual $1-3\cos^2(\theta)$ form of the resonant dipolar spin-exchange.

physics.atom-ph

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Inference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the $\textbf{D}$ecoupled Clip and $\textbf{D}$ynamic s$\textbf{A}$mpling $\textbf{P}$olicy $\textbf{O}$ptimization ($\textbf{DAPO}$) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL.

cs.LG

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery

Understanding the chemical structure from a graphical representation of a molecule is a challenging image caption task that would greatly benefit molecule-centric scientific discovery. Variations in molecular images and caption subtasks pose a significant challenge in both image representation learning and task modeling. Yet, existing methods only focus on a specific caption task that translates a molecular image into its graph structure, i.e., OCSR. In this paper, we propose the Optical Chemical Structure Understanding (OCSU) task, which extends low-level recognition to multilevel understanding and aims to translate chemical structure diagrams into readable strings for both machine and chemist. To facilitate the development of OCSU technology, we explore both OCSR-based and OCSR-free paradigms. We propose DoubleCheck to enhance OCSR performance via attentive feature enhancement for local ambiguous atoms. It can be cascaded with existing SMILES-based molecule understanding methods to achieve OCSU. Meanwhile, Mol-VL is a vision-language model end-to-end optimized for OCSU. We also construct Vis-CheBI20, the first large-scale OCSU dataset. Through comprehensive experiments, we demonstrate the proposed approaches excel at providing chemist-readable caption for chemical structure diagrams, which provide solid baselines for further research. Our code, model, and data are open-sourced at https://github.com/PharMolix/OCSU.

cs.CV

Realization of a doped quantum antiferromagnet with dipolar tunnelings in a Rydberg tweezer array

Doping an antiferromagnetic Mott insulator is central to our understanding of a variety of phenomena in strongly-correlated electrons, including high-temperature superconductors. To describe the competition between tunneling $t$ of hole dopants and antiferromagnetic (AFM) spin interactions $J$, theoretical and numerical studies often focus on the paradigmatic $t$-$J$ model, and the direct analog quantum simulation of this model in the relevant regime of high-particle density has long been sought. Here, we realize a doped quantum antiferromagnet with next-nearest neighbour (NNN) tunnelings $t'$ and hard-core bosonic holes using a Rydberg tweezer platform. We utilize coherent dynamics between three Rydberg levels, encoding spins and holes, to implement a tunable bosonic $t$-$J$-$V$ model allowing us to study previously inaccessible parameter regimes. We observe dynamical phase separation between hole and spin domains for $|t/J|\ll 1$, and demonstrate the formation of repulsively bound hole pairs in a variety of spin backgrounds. The interference between NNN tunnelings $t'$ and perturbative pair tunneling gives rise to light and heavy pairs depending on the sign of $t$. Using the single-site control allows us to study the dynamics of a single hole in 2D square lattice (anti)ferromagnets. The model we implement extends the toolbox of Rydberg tweezer experiments beyond spin-1/2 models to a larger class of $t$-$J$ and spin-$1$ models.

quant-ph

Tomonaga-Luttinger Liquid Behavior in a Rydberg-encoded Spin Chain

Quantum fluctuations can disrupt long-range order in one-dimensional systems, and replace it with the universal paradigm of the Tomonaga-Luttinger liquid (TLL), a critical phase of matter characterized by power-law decaying correlations and linearly dispersing excitations. Using a Rydberg quantum simulator, we study how TLL physics manifests in the low-energy properties of a spin chain, interacting under either the ferromagnetic or the antiferromagnetic dipolar XY Hamiltonian. Following quasi-adiabatic preparation, we directly observe the power-law decay of spin-spin correlations in real-space, allowing us to extract the Luttinger parameter. In the presence of an impurity, the chain exhibits tunable Friedel oscillations of the local magnetization. Moreover, by utilizing a quantum quench, we directly probe the propagation of correlations, which exhibit a light-cone structure related to the linear sound mode of the underlying TLL. Our measurements demonstrate the influence of the long-range dipolar interactions, renormalizing the parameters of TLL with respect to the case of nearest-neighbor interactions. Finally, comparison to numerical simulations exposes the high sensitivity of TLLs to doping and finite-size effects.

quant-ph

Reach Measurement, Optimization and Frequency Capping In Targeted Online Advertising Under k-Anonymity

The growth in the use of online advertising to foster brand awareness over recent years is largely attributable to the ubiquity of social media. One pivotal technology contributing to the success of online brand advertising is frequency capping, a mechanism that enables marketers to control the number of times an ad is shown to a specific user. However, the very foundation of this technology is being scrutinized as the industry gravitates towards advertising solutions that prioritize user privacy. This paper delves into the issue of reach measurement and optimization within the context of $k$-anonymity, a privacy-preserving model gaining traction across major online advertising platforms. We outline how to report reach within this new privacy landscape and demonstrate how probabilistic discounting, a probabilistic adaptation of traditional frequency capping, can be employed to optimize campaign performance. Experiments are performed to assess the trade-off between user privacy and the efficacy of online brand advertising. Notably, we discern a significant dip in performance as long as privacy is introduced, yet this comes with a limited additional cost for advertising platforms to offer their users more privacy.

cs.GT

A Single-Ion Information Engine for Charging Quantum Battery

Information engines produce mechanical work through measurement and adaptive control. For information engines, the principal challenge lies in how to store the generated work for subsequent utilization. Here, we report an experimental demonstration where quantized mechanical motion serves as a quantum battery and gets charged in repeated cycles by a single trapped-ion information engine. This is enabled by a key technological advancement in rapid state discrimination, allowing us to suppress measurement-induced disturbances. Consequently, we were able to obtain a charging efficiency over 50\% of the theoretical limit at the optimal temperature. The experimental results substantiate that this approach can render trapped ions a promising platform for microscopic information engines with potential applications in the future upon scaling up.

quant-ph

BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

Foundation models (FMs) have exhibited remarkable performance across a wide range of downstream tasks in many domains. Nevertheless, general-purpose FMs often face challenges when confronted with domain-specific problems, due to their limited access to the proprietary training data in a particular domain. In biomedicine, there are various biological modalities, such as molecules, proteins, and cells, which are encoded by the language of life and exhibit significant modality gaps with human natural language. In this paper, we introduce BioMedGPT, an open multimodal generative pre-trained transformer (GPT) for biomedicine, to bridge the gap between the language of life and human natural language. BioMedGPT allows users to easily ``communicate'' with diverse biological modalities through free text, which is the first of its kind. BioMedGPT aligns different biological modalities with natural language via a large generative language model, namely, BioMedGPT-LM. We publish BioMedGPT-10B, which unifies the feature spaces of molecules, proteins, and natural language via encoding and alignment. Through fine-tuning, BioMedGPT-10B outperforms or is on par with human and significantly larger general-purpose foundation models on the biomedical QA task. It also demonstrates promising performance in the molecule QA and protein QA tasks, which could greatly accelerate the discovery of new drugs and therapeutic targets. In addition, BioMedGPT-LM-7B is the first large generative language model based on Llama2 in the biomedical domain, therefore is commercial friendly. Both BioMedGPT-10B and BioMedGPT-LM-7B are open-sourced to the research community. In addition, we publish the datasets that are meticulously curated for the alignment of multi-modalities, i.e., PubChemQA and UniProtQA. All the models, codes, and datasets are available at \url{https://github.com/PharMolix/OpenBioMed}.

cs.CE

GNN-Ensemble: Towards Random Decision Graph Neural Networks

Graph Neural Networks (GNNs) have enjoyed wide spread applications in graph-structured data. However, existing graph based applications commonly lack annotated data. GNNs are required to learn latent patterns from a limited amount of training data to perform inferences on a vast amount of test data. The increased complexity of GNNs, as well as a single point of model parameter initialization, usually lead to overfitting and sub-optimal performance. In addition, it is known that GNNs are vulnerable to adversarial attacks. In this paper, we push one step forward on the ensemble learning of GNNs with improved accuracy, generalization, and adversarial robustness. Following the principles of stochastic modeling, we propose a new method called GNN-Ensemble to construct an ensemble of random decision graph neural networks whose capacity can be arbitrarily expanded for improvement in performance. The essence of the method is to build multiple GNNs in randomly selected substructures in the topological space and subfeatures in the feature space, and then combine them for final decision making. These GNNs in different substructure and subfeature spaces generalize their classification in complementary ways. Consequently, their combined classification performance can be improved and overfitting on the training data can be effectively reduced. In the meantime, we show that GNN-Ensemble can significantly improve the adversarial robustness against attacks on GNNs.

cs.LG