arXiv ScienceSearch

arXiv subjects

Pan Yu

Publications and source records attributed to Pan Yu.

3 recordsLinked to original sources

Light-Cone Scaling of In-Circuit Noise in Randomized Measurements

Randomized measurements provide an efficient way to extract physical properties of an unknown quantum state from limited data. On near-term hardware, gate and readout errors bias the reconstructed observables. Here we develop a microscopic description of this bias for locally scrambled shallow circuits. Independent local twirling reduces local implementation noise to stochastic Pauli damping, and a noise event contributes only when it overlaps the Heisenberg evolution of the measured Pauli operator. This gives an activated path-average formula for the noisy Pauli coefficient. In one-dimensional shallow circuits, the activated noise volume grows linearly with the size of a contiguous observable, leading to an exponential damping ratio. We verify this scaling for two-qubit random Clifford and locally scrambled iSWAP circuits with two-qubit Pauli noise, including spatial fluctuations and temporal drift. The scaling supports a small-string calibration protocol that predicts larger string observables without learning the full noisy measurement channel. Our result relates the noise bias of shallow-shadow protocols directly to operator-evolving dynamics.

quant-ph

Global Prompt Refinement with Non-Interfering Attention Masking for One-Shot Federated Learning

Federated Prompt Learning (FPL) enables communication-efficient adaptation by tuning lightweight prompts on top of frozen pre-trained models. Existing FPL methods typically rely on global information, which is only available after the second training round, to facilitate collaboration among client models. Therefore, they are inherently dependent on multi-round communication to fully exhibit their strengths. Moreover, existing one-shot federated learning methods typically focus on fitting seen tasks, but lack cross-task generalization. To bridge this gap, we propose the Global Prompt Refinement with Non-Interfering Attention Masking (GPR-NIAM) method for one-shot FPL. The core idea is to design a masking mechanism that restricts excessive interaction between the original text embeddings and the learnable prompt embeddings. GPR-NIAM achieves this through the collaboration of two key modules. Firstly, the attention isolation module suppresses attention from the learnable prompt tokens to the original text tokens, and reweights the reverse attention which preserves generalization across tasks. Secondly, the cross-silo collaborative refinement module integrates decentralized visual knowledge into a unified base and calibrates the global prompt through multi-source cross-modal knowledge alignment, further mitigating the inconsistency caused by data heterogeneity. Extensive experiments conducted on ten benchmark datasets under two tasks show that GPR-NIAM outperforms eight state-of-the-art methods in both class-level and domain-level generalization.

cs.CV

Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss Optimization

Existing audio-driven facial animation methods face critical challenges, including expression leakage, ineffective subtle expression transfer, and imprecise audio-driven synchronization. We discovered that these issues stem from limitations in motion representation and the lack of fine-grained control over facial expressions. To address these problems, we present Takin-ADA, a novel two-stage approach for real-time audio-driven portrait animation. In the first stage, we introduce a specialized loss function that enhances subtle expression transfer while reducing unwanted expression leakage. The second stage utilizes an advanced audio processing technique to improve lip-sync accuracy. Our method not only generates precise lip movements but also allows flexible control over facial expressions and head motions. Takin-ADA achieves high-resolution (512x512) facial animations at up to 42 FPS on an RTX 4090 GPU, outperforming existing commercial solutions. Extensive experiments demonstrate that our model significantly surpasses previous methods in video quality, facial dynamics realism, and natural head movements, setting a new benchmark in the field of audio-driven facial animation.

cs.CV