arXiv ScienceSearch

arXiv subjects

Yifang Xu

Publications and source records attributed to Yifang Xu.

At least 19 recordsLinked to original sources

Experimental realization of quantum-computing-enhanced sensing

Quantum algorithms offer computational advantages, yet incorporating them into quantum sensing and converting these advantages into enhanced information acquisition remains challenging. Here, we realize quantum-computing-enhanced sensing in a spin-oscillator architecture, where a microwave cavity provides a high-dimensional quantum register and a coupled superconducting qubit serves as the sensor. The unknown signal directly generates the phase oracle operation, establishing a natural physical interface between quantum computation and sensing. We experimentally demonstrate the first Grover search in a bosonic mode and observe quantum amplitude amplification in a Hilbert space spanning more than 128 photons. For the same number of sensing iterations, our protocol extracts more information and resolves more frequency candidates than classical sequential search. Our results establish signal-driven oracles as a route for harnessing oracle-based quantum algorithms and realizing quantum-computing-enhanced sensing.

quant-ph

Quantum Optical Reinforcement Learning via Spectrum-Resolved Hong-Ou-Mandel Interference

Hong-Ou-Mandel (HOM) interference-based optical neural networks can offer complexity advantages on benchmark learning tasks, but conventional readout compresses the coincidence spectrum into a single scalar, limiting its use in complex settings such as continuous-action reinforcement learning. Here we introduce a spectrum-resolved HOM (SR-HOM) architecture that promotes the photons' spectral degrees of freedom to a trainable computational resource and use it to construct a compact optical actor-critic agent. Diagonal spectral responses generate continuous actions, while higher-order spectral correlations provide nonlinear state-action features for value estimation. Across five continuous-control benchmarks, SR-HOM outperforms parameter-matched multilayer-perceptron baselines, including a \(4.4\times\) improvement in sample efficiency and a \(74.0\%\) increase in best 100-episode moving-average return for LunarLanderContinuous-v3. Applied to online calibration of drifted tunable-coupler CZ and iSWAP gates for transmon qubits, simulations show it restores fidelities to \(0.9917\) and \(0.9952\) respectively, exceeding \(99.8\%\) of their drift-free calibrated values.

quant-ph

Robust and optimal control of open quantum systems

Recent advancements in quantum technologies have highlighted the importance of mitigating system imperfections, including parameter uncertainties and decoherence effects, to improve the performance of experimental platforms. However, most of the previous efforts in quantum control are devoted to the realization of arbitrary unitary operations in a closed quantum system. Here, we improve the algorithm that suppresses system imperfections and noises, providing notably enhanced scalability for robust and optimal control of open quantum systems. Through experimental validation in a superconducting quantum circuit, we demonstrate that our approach outperforms its conventional counterpart for closed quantum systems with an ultra-low infidelity of about $0.60\%$, while the complexity of this algorithm exhibits the same scaling, with only a modest increase in the prefactor. This work represents a notable advancement in quantum optimal control techniques, paving the way for realizing quantum-enhanced technologies in practical applications.

quant-ph

Quantum Confocal Microscopy in Fock Space with a 19 dB Metrological Gain

Quantum metrology promises measurement precision beyond classical limits by exploiting large-scale quantum states, yet realizing this advantage faces two fundamental challenges: the deterministic preparation of non-trivial quantum probes and the efficient extraction of metrological information in high-dimensional Hilbert spaces. Here, we introduce quantum confocal microscopy in Fock space that simultaneously resolves both challenges. Drawing a direct analogy between classical wave optics and quantum state evolution in a bosonic mode, we construct a confocal system with two Fock-space lenses. The first lens deterministically focuses a coherent state into a quantum probe with a tightly concentrated photon-number distribution, while the second lens maps the metrological information back to the vacuum state for efficient readout. Using a superconducting circuit QED platform, we prepare focused probe states with mean photon numbers up to ${N} = 500$, achieving a 21.5$\pm$1.1 dB compression of the photon-number uncertainty relative to a coherent state, with a scalable quantum circuit of $\mathcal{O}(1)$ operational depth. We demonstrate a displacement sensitivity scaling as $N^{-0.416}$, approaching the Heisenberg scaling ($N^{-0.5}$), and achieve a record metrological gain of 19.06$\pm$0.13 dB beyond the standard quantum limit. This work establishes quantum confocal microscopy as a scalable and practical framework for quantum-enhanced precision measurement, readily extendable to other bosonic platforms and high-dimensional quantum many-body systems.

quant-ph

Deterministic Generation of Arbitrary Fock States via Resonant Subspace Engineering

Deterministic preparation of high-excitation Fock states is a central challenge in bosonic quantum information, with control complexity that generically explodes as the Hilbert space dimension grows. Here we introduce resonant subspace engineering (RSE), a protocol that analytically confines the infinite-dimensional bosonic dynamics to a two-dimensional invariant subspace spanned by an initial coherent state and the target state. State transfer then reduces to a geodesic rotation on a synthetic Bloch sphere, governed by resonance and phase-matching conditions we derive in closed form. For single Fock states, RSE achieves $O(n^{1/4})$ scaling in both evolution time and gate depth, showing a fundamental improvement over existing deterministic schemes. The construction generalizes to $K$-component superpositions via a $(K{+}1)$-dimensional invariant subspace with full $\mathrm{SU}(K{+}1)$ controllability, requiring only 3-5 iterations of operations for superpositions spanning photon numbers 70--100. RSE provides a scalable and analytically transparent framework for large-scale bosonic state engineering and gate synthesis across single- and multimode platforms.

quant-ph

FaceSnap: Enhanced ID-fidelity Network for Tuning-free Portrait Customization

Benefiting from the significant advancements in text-to-image diffusion models, research in personalized image generation, particularly customized portrait generation, has also made great strides recently. However, existing methods either require time-consuming fine-tuning and lack generalizability or fail to achieve high fidelity in facial details. To address these issues, we propose FaceSnap, a novel method based on Stable Diffusion (SD) that requires only a single reference image and produces extremely consistent results in a single inference stage. This method is plug-and-play and can be easily extended to different SD models. Specifically, we design a new Facial Attribute Mixer that can extract comprehensive fused information from both low-level specific features and high-level abstract features, providing better guidance for image generation. We also introduce a Landmark Predictor that maintains reference identity across landmarks with different poses, providing diverse yet detailed spatial control conditions for image generation. Then we use an ID-preserving module to inject these into the UNet. Experimental results demonstrate that our approach performs remarkably in personalized and customized portrait generation, surpassing other state-of-the-art methods in this domain.

cs.CV

Diff-PC: Identity-preserving and 3D-aware Controllable Diffusion for Zero-shot Portrait Customization

Portrait customization (PC) has recently garnered significant attention due to its potential applications. However, existing PC methods lack precise identity (ID) preservation and face control. To address these tissues, we propose Diff-PC, a diffusion-based framework for zero-shot PC, which generates realistic portraits with high ID fidelity, specified facial attributes, and diverse backgrounds. Specifically, our approach employs the 3D face predictor to reconstruct the 3D-aware facial priors encompassing the reference ID, target expressions, and poses. To capture fine-grained face details, we design ID-Encoder that fuses local and global facial features. Subsequently, we devise ID-Ctrl using the 3D face to guide the alignment of ID features. We further introduce ID-Injector to enhance ID fidelity and facial controllability. Finally, training on our collected ID-centric dataset improves face similarity and text-to-image (T2I) alignment. Extensive experiments demonstrate that Diff-PC surpasses state-of-the-art methods in ID preservation, facial control, and T2I consistency. Furthermore, our method is compatible with multi-style foundation models.

cs.CV

Principles of Optics in the Fock Space: Scalable Manipulation of Giant Quantum States

The manipulation of distinct degrees of freedom of photons plays a critical role in both classical and quantum information processing. While the principles of wave optics provide elegant and scalable control over classical light in spatial and temporal domains, engineering quantum states in Fock space has been largely restricted to few-photon regimes, hindered by the computational and experimental challenges of large Hilbert spaces. Here, we introduce ``Fock-space optics", establishing a conceptual framework of wave propagation in the quantum domain by treating photon number as a synthetic dimension. Using a superconducting microwave resonator, we experimentally demonstrate Fock-space analogues of optical propagation, refraction, lensing, dispersion, and interference with up to 180 photons. These results establish a fundamental correspondence between Schr\"{o}dinger evolution in a single bosonic mode and classical paraxial wave propagation. By mapping intuitive optical concepts onto high-dimensional quantum state engineering, our work opens a path toward scalable control of large-scale quantum systems with thousands of photons and advanced bosonic information processing.

quant-ph

Scalable Generation of Macroscopic Fock States Exceeding 10,000 Photons

The scalable preparation of bosonic quantum states with macroscopic excitations poses a fundamental challenge in quantum technologies, limited by control complexity and photon-loss rates that severely constrain prior theoretical and experimental efforts to merely dozens of excitations per mode. Here, based on the duality of the quantum state evolution in Fock state space and the optical wave-function propagation in a waveguide array, we introduce a Kerr-engineered multi-lens protocol in a single bosonic mode to deterministically generate Fock states exceeding $10,000$ photons. By optimizing phase and displacement operations across lens groups, our approach compensates for non-paraxial aberrations, achieving fidelities above $73\%$ in numerical simulations for photon numbers up to $N=100,000$. Counterintuitively, the protocol's execution time scales as $N^{-1/2}$ with the target photon number $N$, exhibiting robustness against the photon loss. Our framework enables exploration of quantum-to-classical transitions of giant Fock states, paving the way for advanced quantum metrology with significant quantum gains, and error-corrected quantum information processing in high-dimensional Hilbert spaces.

quant-ph

High-performance quantum interconnect between bosonic modules beyond transmission loss constraints

Distributed quantum computing architectures require high-performance quantum interconnects between quantum information processing units, while previous implementations have been fundamentally limited by transmission line losses. Here, we demonstrate a low-loss interconnect between two superconducting modules using an aluminum coaxial cable, achieving a bus mode quality factor of 1.7e6. By employing SNAIL as couplers, we realize inter-modular state transfer in 0.8 {\mu}s via a three-wave mixing process. The state transfer fidelity reaches 98.2% for quantum states encoded in the first two energy levels, achieving a Bell state fidelity of 92.5%. Furthermore, we show the capability to transfer high-dimensional states by successfully transmitting binomially encoded logical states. Systematic characterization reveals that performance constraints have shifted from transmission line losses (contributing merely 0.2% infidelity) to module-channel interface effects and local Kerr nonlinearities. Our work advances the realization of quantum interconnects approaching fundamental capacity limits, paving the way for scalable distributed quantum computing and efficient quantum communications.

quant-ph

Giant-atom quantum acoustodynamics in hybrid superconducting-phononic integrated circuits

We demonstrate a giant atom by coupling a superconducting transmon qubit to a lithium niobate phononic waveguide at two points separated by about 600 acoustic wavelengths, with a propagation delay of 125 ns. The giant atom yields non-Markovian relaxation dynamics characterized by phonon backflow and a frequency-dependent effective decay rate varying four-fold over merely 4 MHz, corresponding to a Purcell factor exceeding 40. Exploiting this frequency-dependent dissipation, we prepare quantum superposition states with high purity. Our results establish phononic integrated circuits as a versatile platform for giant-atom physics, providing highly tunable quantum devices for advanced quantum information processing.

quant-ph

HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion

Recent advancements in diffusion-based technologies have made significant strides, particularly in identity-preserved portrait generation (IPG). However, when using multiple reference images from the same ID, existing methods typically produce lower-fidelity portraits and struggle to customize face attributes precisely. To address these issues, this paper presents HiFi-Portrait, a high-fidelity method for zero-shot portrait generation. Specifically, we first introduce the face refiner and landmark generator to obtain fine-grained multi-face features and 3D-aware face landmarks. The landmarks include the reference ID and the target attributes. Then, we design HiFi-Net to fuse multi-face features and align them with landmarks, which improves ID fidelity and face control. In addition, we devise an automated pipeline to construct an ID-based dataset for training HiFi-Portrait. Extensive experimental results demonstrate that our method surpasses the SOTA approaches in face similarity and controllability. Furthermore, our method is also compatible with previous SDXL-based works.

cs.CV

WAM-Diff: A Masked Diffusion VLA Framework with MoE and Online Reinforcement Learning for Autonomous Driving

End-to-end autonomous driving systems based on vision-language-action (VLA) models integrate multimodal sensor inputs and language instructions to generate planning and control signals. While autoregressive large language models and continuous diffusion policies are prevalent, the potential of discrete masked diffusion for trajectory generation remains largely unexplored. This paper presents WAM-Diff, a VLA framework that employs masked diffusion to iteratively refine a discrete sequence representing future ego-trajectories. Our approach features three key innovations: a systematic adaptation of masked diffusion for autonomous driving that supports flexible, non-causal decoding orders; scalable model capacity via a sparse MoE architecture trained jointly on motion prediction and driving-oriented visual question answering (VQA); and online reinforcement learning using Group Sequence Policy Optimization (GSPO) to optimize sequence-level driving rewards. Remarkably, our model achieves 91.0 PDMS on NAVSIM-v1 and 89.7 EPDMS on NAVSIM-v2, demonstrating the effectiveness of masked diffusion for autonomous driving. The approach provides a promising alternative to autoregressive and diffusion-based policies, supporting scenario-aware decoding strategies for trajectory generation. The code for this paper will be released publicly at: https://github.com/fudan-generative-vision/WAM-Diff

cs.RO

WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving

We introduce WAM-Flow, a vision-language-action (VLA) model that casts ego-trajectory planning as discrete flow matching over a structured token space. In contrast to autoregressive decoders, WAM-Flow performs fully parallel, bidirectional denoising, enabling coarse-to-fine refinement with a tunable compute-accuracy trade-off. Specifically, the approach combines a metric-aligned numerical tokenizer that preserves scalar geometry via triplet-margin learning, a geometry-aware flow objective and a simulator-guided GRPO alignment that integrates safety, ego progress, and comfort rewards while retaining parallel generation. A multi-stage adaptation converts a pre-trained auto-regressive backbone (Janus-1.5B) from causal decoding to non-causal flow model and strengthens road-scene competence through continued multimodal pretraining. Thanks to the inherent nature of consistency model training and parallel decoding inference, WAM-Flow achieves superior closed-loop performance against autoregressive and diffusion-based VLA baselines, with 1-step inference attaining 89.1 PDMS and 5-step inference reaching 90.3 PDMS on NAVSIM v1 benchmark. These results establish discrete flow matching as a new promising paradigm for end-to-end autonomous driving. The code will be publicly available soon.

cs.RO

Circuit Quantum Acoustodynamics in a Scalable Phononic Integrated Circuit Architecture

Previous demonstrations of quantum acoustic systems have been limited to isolated devices, with limited capability to route phonons and interconnect multi-port acoustic elements for further extension. Here, we demonstrate a scalable architecture for circuit quantum acoustodynamics (cQAD) by integrating superconducting qubits with suspension-free phononic integrated circuits (PnICs). Coherent coupling between tunable transmon qubits and waveguide-integrated phononic cavities, including Fabry-Perot cavities via monolithic integration and microring cavities via flip-chip assembly, has been achieved, producing a pronounced enhancement of phonon emission with a Purcell factor of ~19. These devices represent elementary building blocks for scalable phononic circuits, establishing the foundation for phonon-based quantum information processors and the testbed for novel quantum acoustic phenomena.

quant-ph

High-Fidelity Controlled-Phase Gate for Binomial Codes via Geometric Phase Engineering

High-fidelity two-logical-qubit gates are essential for realizing fault-tolerant quantum computation with bosonic codes, yet experimentally reported fidelities have rarely exceeded 90\%. Here, we propose a geometric phase engineering approach for implementing controlled-phase gates for binomially encoded logical qubits. This method leverages the structural simplicity of geometric drives to reduce the numerical optimization dimensionality while fully incorporating system nonlinearities, enabling fast and high-fidelity logical operations. As an example, we experimentally demonstrate a process fidelity of 97.4$\pm$0.8\% for a controlled-Z gate between two binomial codes, surpassing all previously reported two-logical-qubit gates in bosonic codes. This work demonstrates that geometric phase engineering provides an effective and experimentally feasible route to fast, high-fidelity logical operations in bosonic quantum processors.

quant-ph

Extending coherence time beyond break-even point using only drives and dissipation

Quantum error correction (QEC) aims to mitigate the loss of quantum information to the environment, which is a critical requirement for practical quantum computing. Existing QEC implementations heavily rely on measurement-based feedback, however, constraints on readout fidelity, hardware latency, and system complexity often limit both performance and scalability. Autonomous QEC (AQEC) seeks to overcome these obstacles by stabilizing logical codewords using introduced drives that provide coherent control and engineered dissipation. Here, we propose an AQEC protocol, derived from quantum channel simulation, that is applicable to arbitrary error-correcting codes. As a demonstration, we implement the protocol using a binomial code encoded in a long-lived bosonic mode (lifetime > 1ms), and extend the logical qubit coherence time to 1.04 times that of the best physical qubit in the system. This is the first experimental realization of an AQEC-protected bosonic logical qubit beyond the break-even point, proving that coherence time can indeed be extended by introducing only drives and dissipation. Our results highlight the performance and scalability potential of AQEC, marking an important step toward large-scale, universal quantum computing.

quant-ph

ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images

Recent advances in Multimodal Large Language Models (MLLMs) have introduced a paradigm shift for Image Quality Assessment (IQA) from unexplainable image quality scoring to explainable IQA, demonstrating practical applications like quality control and optimization guidance. However, current explainable IQA methods not only inadequately use the same distortion criteria to evaluate both User-Generated Content (UGC) and AI-Generated Content (AIGC) images, but also lack detailed quality analysis for monitoring image quality and guiding image restoration. In this study, we establish the first large-scale Visual Distortion Assessment Instruction Tuning Dataset for UGC images, termed ViDA-UGC, which comprises 11K images with fine-grained quality grounding, detailed quality perception, and reasoning quality description data. This dataset is constructed through a distortion-oriented pipeline, which involves human subject annotation and a Chain-of-Thought (CoT) assessment framework. This framework guides GPT-4o to generate quality descriptions by identifying and analyzing UGC distortions, which helps capturing rich low-level visual features that inherently correlate with distortion patterns. Moreover, we carefully select 476 images with corresponding 6,149 question answer pairs from ViDA-UGC and invite a professional team to ensure the accuracy and quality of GPT-generated information. The selected and revised data further contribute to the first UGC distortion assessment benchmark, termed ViDA-UGC-Bench. Experimental results demonstrate the effectiveness of the ViDA-UGC and CoT framework for consistently enhancing various image quality analysis abilities across multiple base MLLMs on ViDA-UGC-Bench and Q-Bench, even surpassing GPT-4o.

cs.CV