arXiv ScienceSearch

arXiv subjects

Yuan Yao

Publications and source records attributed to Yuan Yao.

At least 19 recordsLinked to original sources

Low-cost algorithm-to-execution framework for surface-code quantum computing

The execution of useful quantum algorithms on fault-tolerant processors requires more than a mapping from logical gates to encoded operations: the spatial organization, non-Clifford resource supply, and execution schedule must also be determined while keeping physical overhead within practical limits. Although the theoretical hierarchy from logical circuits to fault-tolerant operations is well established, these implementation choices are often specified and optimized separately. Here we develop a low-cost algorithm-to-execution framework for surface-code quantum computing. From hierarchical algorithm descriptions, it constructs dependency-preserving logical schedules and an executable workload capturing logical interactions, operation parallelism, and time-resolved non-Clifford demand, thereby linking logical computation to surface-code organization, resource-state preparation, and fault-tolerant execution in a traceable workflow. We apply the framework to twenty benchmark circuits across seven algorithm families and a hierarchically composed application-scale elliptic-curve discrete-logarithm workload. Physical costs vary substantially even for circuits with similar logical resource counts. Under our direct-rotation calibration, non-Clifford implementation selection reduces space-time volume by up to 241.5 times versus an all-synthesis baseline for the QAOA amplitude-amplification workload. Circuit-specific surface-code layouts reduce routed-latency estimates for all twenty benchmarks; thirteen also reduce space-time volume because communication savings outweigh added spatial overhead. These results show that low-cost fault-tolerant execution depends on computation scheduling and organization, not aggregate logical resource counts alone.

quant-ph

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms: retrieval banks, learned compressors, recurrent states must decide what to keep from the past before knowing what a future decision will require. This was motivated by the assumption that minute-scale history is too large to process directly, which modern VLM backbones no longer make true. In this work, we introduce SimpleMemVLA, a VLA without a dedicated memory module. It keeps the sampled history intact and passes it to the backbone in the timestamped video format the backbone was pretrained to process; the hidden states of a generated sub-task then form the only channel from history to a standard flow-matching action head. Since consecutive decisions share most of their history, prefilling the shared prefix during action execution keeps latency close to a single-frame VLA. SimpleMemVLA sets a new state of the art on four memory benchmarks without cost on general-purpose control. Holding the backbone and training setup fixed, it outperforms retrieval, compression and recurrent-state mechanisms by a wide margin, and causal interventions confirm that the policy genuinely reads its history. Code available at https://github.com/wadeKeith/SimpleMemVLA

cs.CV

Large Exchange Magnetostriction in a Kagome Antiferromagnet at Room Temperature

The pursuit of high-performance magnetostrictive materials is crucial for advancing technologies in sensing, actuation, and microelectromechanical systems. Although pronounced magnetostrictive effects have been observed in a few ferromagnets, a systematic exploration of magnetostriction across a broader range of antiferromagnets remains limited. Here, we report the observation of large magnetostriction in the kagome antiferromagnet YMn6Sn6. Under a magnetic field, the system undergoes a field-induced evolution from a helical magnetic ground state toward a collinear field-polarized state, accompanied by a large anisotropic lattice strain and a substantial volume magnetostriction exceeding 400 ppm at room temperature. Remarkably, the magnetostrictive response remains nearly fully reversible up to 9 T with nearly hysteresis-free behavior, effectively minimizing the energy dissipation commonly associated with domain-wall pinning. Combined experimental measurements and theoretical calculations reveal that the large, nearly hysteresis-free magnetostriction originates from the competition between intralayer ferromagnetic and interlayer exchange interactions. This exchange-driven magnetoelastic coupling further gives rise to a strongly direction-dependent lattice distortion pathway. This work establishes kagome helimagnets as a tunable platform for low-dissipation magnetoelastic functionalities and responsive magnetomechanical applications.

cond-mat.mtrl-sci

Mobile Backscatter Communication for the Battery-less Internet of Things

We enable backscatter communication in the battery-less mobile Internet of Things (IoT). Backscatter communication is extensively studied in static settings. Existing designs are, however, fundamentally mismatched with mobility and time-varying energy patterns. Channel conditions rapidly fluctuate, impacting the achievable data rates and thus transmission costs. Energy availability varies unpredictably, possibly forcing devices to remain quiescent to recharge energy buffers. The two issues compound each other: while recharging, a battery-less mobile IoT device may miss more favorable channel conditions. We design a lightweight decision system that dynamically determines when to transmit by checking short-term trends in signal strength, while using Non-volatile Memory (NVM) to retain packets in unfavorable channel conditions and across energy failures. Using a prototype we built and real-world mobility and power traces, we compare our design against a rate-adaptive baseline that only considers the instantaneous channel conditions. Experimental results show that our system improves throughput by up to 5.16x while reducing transmission energy consumption by up to 47.3%, with only 0.23% - 7.3% additional energy overhead.

cs.NI

Symmetry-Enforced Topological Structures in Quantum Phase Diagrams

We study the topological structure of the quantum phase diagram of gapped systems by identifying the noncontractibility of loops, so-called $S^1$-families, within the gapped phase diagram in which many-body Hamiltonians can have nontrivial ground-state degeneracy. We manifest the role of symmetries in such $S^1$-family classifications by the exotic symmetry interplay: (i) nontrivial mixed anomalies, (ii) semi-direct product relation between the spontaneously broken and the unbroken symmetries, and (iii) symmetry with noninvertible operators. We find that such structures lead to the novel $S^1$-family classifications inaccessible by earlier classifications based on ``independent'' symmetries. Furthermore, we construct lattice realizations of these $S^1$-families and explicitly demonstrate their novel algebraic structures.

cond-mat.str-el

Zero-point theorems in quantum many-body physics

We propose several zero-point type arguments based on the inevitable zero point(s) of a spectral gap in the quantum spin system phase diagrams in various dimensions. We consider multi-parameter families of Hamiltonian extending the conventional zero-point theorem that includes only one parameter. Analogously to the zero-point theorem, we only impose model-independent transformation relations along the parameter boundary, rather than specifying any low-energy dynamics or response. We further give a series of conjectures, which generalize our statements in a uniform way. Our results give powerful and universal model-independent constraints on the possible relevant operators for critical phenomena in quantum spin models in arbitrary high dimensions.

cond-mat.str-el

Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution

Effective research ideation requires moving beyond a static understanding of prior work to trace how research problems and solutions evolve across the literature. Existing methods either treat papers as unstructured context or model scholarly evolution as isolated citation chains, overlooking interactions among research trajectories. We propose Tree-of-Ideas (ToI), a two-stage framework. EvoTrace reconstructs branching scholarly trajectories from citations, tracking evolving methods, resolved problems, and gaps. EvoAgent then reasons across trajectories to identify convergent problems and complementary solutions, generating grounded research ideas. Across six AI research topics, ToI achieves the highest score among automatic methods (6.27 vs. 5.36 for the strongest baseline on a 10-point scale), with strong Novelty (6.36) and Groundedness (7.00). Also, its score approaches that of human-paper references (6.29), demonstrating the value of cross-path evolutionary reasoning.

cs.AI

Searching for $J$-holomorphic curves via machine: first steps

We assemble numerical algorithms to search for $J$-holomorphic curves in symplectic manifolds. Each algorithm employs several different numerical techniques, each technique addressing a different aspect of the geometric problem. We separately consider both classical Fourier expansion and deep neural networks in our algorithms and compare their performance. Our algorithms take as input a smooth curve in a given homology class and search for a $J$-holomorphic curve in the same homology class. We first verify we can produce explicitly known holomorphic curves in complex manifolds, for example the Weierstrass $\wp$ function on the torus and curves in $S^2\times S^2$ with the standard complex structure. Then we search for $J$-holomorphic curves in $S^2\times S^2$ with non-integrable almost complex structures: essentially we start with a known holomorphic curve in an integrable almost complex structure $J_0$, deform $J_0$ to a nearby nonintegrable almost complex structure $J_\epsilon$, and use our methods to find the nearby $J_\epsilon$-holomorphic curve.

math.SG

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM-Robot on the day of its release. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot. On Cosmos3-Nano-Policy-DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation-dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control-time Roofline, which distinguishes inference-bound from environment-bound control; the measured pi0.5 points on four LIBERO suites are environment-bound while Cosmos3 stays inference-bound. Code and benchmarks: https://github.com/mingti-org/phyai.

cs.AI

Binary X-rays of doubly stochastic matrices

The X-ray of a permutation is a sequence of sums along each diagonal of the associated permutation matrix. They satisfy certain necessary constraints on distribution of the values, which are conjectured to be sufficient when the sequence is binary. By re-expressing the constraints in a form that allows for real-valued relaxations, we prove that these binary sequences are always X-rays of doubly stochastic matrices.

math.CO

DICE: Detailed Inter-Chiplet End-to-End PHY Modeling for Accurate Chiplet Simulation

Scaling monolithic multicores is increasingly constrained by power/thermal limits, yield, and rising manufacturing and testing costs. Chiplet designs address these challenges by partitioning large dies into smaller parts (typically multiple core-complex dies and an I/O die) linked via high-bandwidth physical fabrics (PHY). As bandwidth and wiring density scale, however, these short-reach links are pushed closer to their signal-integrity limits, increasing susceptibility to noise, crosstalk, and channel loss, motivating stronger link-level reliability mechanisms such as forward error correction (FEC). Despite this trend, state-of-the-art simulation infrastructures often approximate inter-chiplet links using oversimplified, fixed-latency models. Such abstractions overlook the inherently dynamic, runtime-dependent behavior of the PHY -- including channel conditions (e.g., signal-to-noise ratio shifts, signal crosstalk, clock jitter), iterative decoder convergence and packet retransmissions, and application dynamics (e.g., LLC-misses that travel across chiplet boundaries) -- all of which are hard to determine offline. We show that neglecting these effects distorts inter-chiplet packet-level timing and high-level performance metrics such as IPC, leading to off-trend simulation results. We present DICE, an in-simulation, runtime PHY modeling in gem5 that captures the end-to-end inter-chiplet datapath, including QC-LDPC encoding/decoding, PAM4 modulation, lossy-channel transmission, LLR-based demodulation, adaptive packet re-sending, and PHY-level flow control between chiplets.

cs.AR

Gaplessness indicator by topologically trivial twisting operators

We propose several general necessary conditions for quantum many-body system in one dimension respecting U(1) symmetry to be gapped. We show that the ground-state expectation value of topologically trivial twisting operators must approach unity in the thermodynamic limit with a certain finite-size scaling. Equivalently, its violation can indicate gaplessness of U(1)-symmetric Hamiltonians. The topological triviality of such a twisting operator enables us to derive infinitely many other gaplessness indicators by static structure factor to any order in real experiments, which are impossible to obtain by earlier topologically nontrivial twisting operators. We also apply analytic and numerical calculations to test the efficiency and consistency of our results.

cond-mat.str-el

Twisting-operator approach to identifying gaplessness in SU($N$) fermionic systems

We propose a general necessary condition for a spinful fermion chain with SU(2) spin-rotation symmetry to be gapped. Specifically, we prove that the expectation value of a properly defined fermionic twisting operator asymptotically approaches unity in any gapped phase with finite ground-state degeneracy, with finite-size corrections bounded by $\mathscr{O}(1/L)$. Consequently, a non-unity value in the thermodynamic limit provides a sufficient criterion for identifying gapless fermion chains. We confirm this criterion using the $s$-wave Bardeen-Cooper-Schrieffer (BCS) Hamiltonian and determinant quantum Monte Carlo (DQMC) simulations of interacting Hubbard models. We further extend the twisting-operator approach to SU($N$)-symmetric fermionic systems, where the gapped ground-state sector must be SU($N$)-singlet and $\langle \hat{\mathcal{U}}^M\rangle=1+\mathscr{O}(1/L)$ by a fermionic twisting operator $\hat{\mathcal{U}}$ with a suitably chosen integer $M$.

cond-mat.str-el

Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes

Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly from bug reports remains underconstrained. Bug reproduction tests (BRTs) help close this gap by turning a bug report into an executable, bug-specific signal that can guide repair and validate candidate patches. Existing work has therefore studied BRT generation as a core subproblem in APR and mainly evaluates a generated BRT using the fail-to-pass (F->P) criterion, which requires the test to fail on the buggy code but pass on the golden fix. We show that F->P alone is insufficient when the goal of a BRT is to improve downstream repair. In particular, some F->P BRTs are lax, reproducing the observed symptom yet still admitting plausible-but-incorrect patches. We formalize this missing quality dimension by separating F->P BRTs into rigorous and lax ones, and show empirically that only the former consistently improve repair success. We further find that co-generation introduces test--fix error coupling, where the in-trajectory fail-to-pass (F->P) check can pass even when both the generated patch and generated test are wrong. Based on these findings, we propose CoHarden, a co-generation framework that uses the Lax signal as an in-loop convergence criterion. CoHarden first generates a test before any fix, then iteratively hardens the test and fix against surviving mutation patches until the generated test no longer admits Lax regressions. Experiments show that CoHarden reaches 69.4% Resolved and 78.9% F->P on SWE-bench Verified, outperforming the strongest fix-only and cogeneration baselines by +9.6 and +7.9 percentage points in Resolved, respectively, with consistent gains across LLM backbones and benchmarks.

cs.SE

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors

Understanding long videos with multimodal large language models (MLLMs) requires selecting a compact set of frames from thousands of candidates, yet identifying the right frames seemingly requires understanding the video first. We resolve this circular dependency with a simple observation: cross-modal attention at validation-selected extraction layers in MLLMs already provides query-relevant frame evidence without requiring autoregressive generation. We exploit this property to build DAFS (Dynamic Attention-based Budget-aware Frame Selection), a training-free frame selector. A lightweight MLLM selector, even with only 2B parameters, can extract frame-level evidence by converting selected-layer attention into relevance scores through query-conditioned aggregation. This enables cross-frame comparison without autoregressive decoding. To handle the selector's own context constraint, we formulate the joint allocation of candidate pool size and per-frame token budget as a discrete optimization problem solved by dynamic programming. Under a 32-frame budget, our selector improves over uniform sampling by up to 6.4 points on Video-MME and outperforms prior training-based selectors under matched frame budgets, while generalizing across selector and answerer backbones, and across tasks, without retraining.

cs.CV

Quantum-classical crossover in fault-tolerant quantum dynamics simulation

While quantum computers promise to solve classically intractable problems, identifying the point at which fault-tolerant quantum computation outperforms the best classical algorithms for practical applications remains an outstanding challenge. Here we establish a concrete quantum-classical crossover for quantum many-body dynamics under realistic hardware conditions. We introduce a scalable fault-tolerant framework that combines coherent observable estimation with a space-time-efficient implementation of non-Clifford rotations, suppressing the residual logical errors that limit existing partially fault-tolerant approaches. A benchmark against state-of-the-art tensor-network and variational Monte Carlo algorithms reveals a concrete crossover for mixed-field Ising dynamics at modest system sizes. For a physical error rate of $p=10^{-3}$, fault-tolerant simulation requires approximately 2 hours and $3.7 \times 10^5$ physical qubits for a 100-site 1D system, whereas tensor network approaches would require about 100 years. For 2D models, where rapid entanglement growth limits the classical evolution time, we project quantum runtimes within minutes. A physical error rate of $p=10^{-4}$ leads to at least an order of magnitude reduction in qubit count ($3.1 \times 10^4$ physical qubits) and runtime (minutes for 1D and seconds for 2D). The reduction in quantum runtime arises from our improved rotation-state injection and co-design of quantum error correction and observable-estimation protocols, which jointly suppress logical-error accumulation and reduce sampling overhead. Our results establish a scalable route towards practical quantum advantage and identify quantitative engineering targets for future fault-tolerant architectures.

quant-ph

Fixed Point Floer Cohomology of Dehn Twists I: Splitting Formulas

This paper is the first in a series following on our earlier work [arXiv:2205.14516, arXiv:2307.08180] studying the pair-of-pants product on fixed point Floer cohomology. In [arXiv:2205.14516, arXiv:2307.08180] we fully computed this product for Dehn twists on surfaces of genus greater or equal to 2, and used it to compute a version of the (small) "quantum cohomology" for nodal curves. In the present work, we develop tools for computing the fixed point Floer cohomology and the associated product in the case of Dehn twists in all higher dimensions: for iterated Dehn twists around a Lagrangian sphere in a Liouville domain, we show that the product and differential on the fixed point Floer cohomology split into local and Morse-theoretic contributions on the level of cochains, using some new confinement results for J-holomorphic curves. The local contributions are expected to recover a finite sector of the homology of (twisted) loop spaces of $S^n$ along with an associated Chas-Sullivan product, which we will examine in detail in future work. We also discuss some immediate applications and curiosities for future work.

math.SG

Who Needs DRAM? We Have Fiber

The rising pressure on DRAM availability and contract pricing reflects generative AI's massive high-performance memory requirements. This pressure is heavily compounded by hyperscale data center expansion, which now consumes a significant portion of global DRAM output. In this work, we propose a new architecture: Fiber Memory, which reimagines the role of optical fiber in a hyperscale data center, deploying it as an active, recirculating delay-line memory for immutable data, such as large language model weights. We present a data-parallel optical broadcast delay-line memory architecture that accounts for fiber's physical realities. By incorporating space-division multiplexed multi-core fibers, passive optical tap-and-amplify interfaces, co-packaged optics, and regional all-optical regeneration, our case study evaluation suggests that Fiber Memory can eliminate redundant weight storage across 10,000 AI accelerators and reduce weight-delivery energy by over 70% compared to traditional HBM3e configurations.

cs.AR