arXiv ScienceSearch

arXiv subjects

Qiang Zhou

Publications and source records attributed to Qiang Zhou.

At least 19 recordsLinked to original sources

Telecom-Integrated Photonic Memory Operating Near the Mechanical Ground State

Scalable quantum networks require quantum memories that are chip-integrated, telecom-band compatible, and capable of flexible retrieval. Nanofabricated mechanical resonators meet these criteria. They offer independent tunability of optical and mechanical modes, long-lived phonon states, and design flexibility beyond atomic systems, making them strong candidates for practical integrated quantum memory. Here, we demonstrate an on-chip, absorptive optomechanical memory for telecom-band photons, based on optomechanically induced transparency (OMIT) and operating near the mechanical ground state. The device stores telecom-band photons, demonstrating compatibility with external photon sources at the few-photon level, while enabling on-demand retrieval. By placing the device in a dilution refrigerator at 20 mK and tailoring the control field to suppress optical heating, we achieve a remarkably low phonon occupancy of just 0.32 during the storage process. Our results lay the groundwork for scalable, phonon-based quantum memory devices and open new avenues for integrating mechanical systems into practical quantum network architectures.

quant-ph

Observation of Hong-Ou-Mandel interference between photon and polariton

Light-matter interactions underlie many quantum technologies, yet whether quasiparticles formed from such interactions preserve the full quantum state of light remains unresolved. Surface plasmon polaritons (SPPs), a class of polaritons formed by interacting photons with free-electron oscillations at metal-dielectric interfaces, are prime candidates to explore this question. Here we demonstrate quantum interference between single photons and SPPs using an Au-SiN$_{\mathrm{x}}$ integrated photonic-plasmonic device. Our results reveal that SPPs retain the indistinguishability of their excitation photons, establishing SPP as a viable quantum information carrier and opening a potential route toward photonic-plasmonic quantum circuitry.

quant-ph

Quantum-interference metrology of dissipative Kerr solitons

Dissipative Kerr solitons in optical microresonators underpin chip-scale frequency combs with applications ranging from coherent telecommunications to precision spectroscopy. Yet the characterization of their intrinsic femtosecond temporal structure remains challenging, as the low pulse energy and broad spectral bandwidth necessitate optical amplification and careful dispersion compensation in conventional ultrafast diagnostics, both of which can significantly distort the waveform. Here we demonstrate a quantum-interference metrology of microcomb solitons based on Hong-Ou-Mandel interference. By attenuating the soliton stream to the single-photon level and measuring fourth-order interference, we directly retrieve near transform-limited pulse durations without amplification or dispersion management, remaining accurate even after propagation through 25 km of standard fiber. The same interferogram also provides direct access to the temporal separations in multi-soliton states by converting inter-soliton separations into additional interference dips at corresponding delays, enabling sub-picosecond characterization of their intracavity temporal structure. This quantum-inspired paradigm introduces a fundamentally new metrological approach that is immune to amplification and dispersion distortions, offering a powerful tool for the characterization of complex soliton physics.

quant-ph

MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends

Most agent-memory benchmarks test post-hoc recall, whereas MemoryArena evaluates whether memory supports interdependent, multi-session task completion. We compare MemoryLake, a structured multi-track memory backend, with Mem0, text-embedding-3-small vector RAG, and a long-context control across all five MemoryArena domains. The systems share the same agent framework, requested gpt-5-mini model alias, task samples, and scoring code; the memory integration is the intentionally changed component. Because each backend bundles write, retrieval, consolidation, budgeting, and prompt-assembly choices, the study is a matched system-level comparison, not a representation-only ablation or a cost-matched experiment. On the shared evaluation sets, MemoryLake has the highest observed success rate (SR) in mathematics (9/40), physics (12/20), and progressive retrieval (4/20). Every system has zero SR in travel planning, and web shopping yields a single bundle-level success (long context, 1/150); MemoryLake ranks third on both the travel soft process score and shopping step match. Following MemoryArena's suite-level convention, a post-hoc equal-weight average over the five SRs is 20.5% for MemoryLake versus 13.6% for the best comparator. These are point estimates: sample sizes are modest, confidence intervals overlap, and we do not report paired significance tests. A separate MemoryLake-only run over all 221 progressive queries yields a failure-counted SR of 26.7% (59/221) and is not a baseline comparison. The results support a workload-dependent view of memory backends and an observed lead among the four evaluated systems on the shared sets; they do not establish benchmark-wide state of the art or a causal advantage of representation structure.

cs.AI

Entanglement-based quantum key distribution with data in hollow-core fiber

The coexistence of quantum information and classical signals in a single fiber is essential for future quantum networks that leverage the well-established optical fiber infrastructure. Although multiplexing technologies can separate quantum and classical signals, pure silica core fibers (PSCFs) remain fundamentally limited by the high nonlinearity, which generates substantial Raman scattering and four-wave mixing noise. Hollow-core fibers (HCFs), guiding light predominantly in air, offer an attractive solution with intrinsically ultra-low nonlinearity and strongly suppressed nonlinear noise. In this work, we demonstrate the entanglement-based key coexisting with data over an 18-km HCF link. We achieve time-encoded high-dimensional quantum key distribution (HD-QKD) carrying 0 dBm of bidirectional received power, corresponding to a theoretical data capacity of up to 2.3 Tbps. During 24 hours of continuous operation, an average secret key rate (SKR) of 10.56 kbps is obtained. Theoretical analysis further predicts SKRs above 135 kbps over transmission distances exceeding 200 km using state-of-the-art low-loss HCFs. These results show significantly improved performance compared with PSCF-based systems and highlight the potential of HCFs for scalable quantum-classical coexistence compatible with the architectures of established fiber-optic networks.

quant-ph

Quantum teleportation over a field-deployed hollow-core fibre network

When a photon and one member of an entangled photon pair are jointly projected onto a Bell-state measurement (BSM), the quantum state of the photon can be transferred to the distant partner of the pair without physically transmitting this information carrier. In real-world deployment, however, teleportation performance is fundamentally bottlenecked by quantum channel impairments, such as loss, noise, and fluctuations, which induce severe decoherence and degrade fidelity. This vulnerability is further exacerbated in scenarios with intense classical data traffic or background light. Realizing scalable quantum networks, therefore, hinges on developing advanced channel architectures capable of supporting both high-fidelity quantum operations and high-capacity classical communications within a shared infrastructure. Towards this end, hollow core fibre (HCF) offers a promising quantum channel resource by combining free-space-like weak light-matter interaction with the stability of fibre-based systems. Here, utilizing a field-deployed metropolitan HCF network spanning three spatially separated nodes in Chengdu, we achieve quantum teleportation with an intermediate BSM under co-propagating classical traffic. Crucially, the HCF links preserve the long-term indistinguishability of photonic qubits without active stabilization, and exhibit a Raman noise approximately three orders of magnitude lower than that of standard solid-core counterparts. This noise suppression enables robust quantum teleportation even alongside classical launch powers up to 160 mW. Our findings establish a classical-data-compatible framework for quantum networking over deployed fibre infrastructure and offer a wavelength-agnostic, plug-and-play, and free-running pathway toward the quantum internet.

quant-ph

Quantum Teleportation toward the Quantum Internet: A Concise Review

Quantum networks play a pivotal role in quantum information science, which not only provide a secure communication platform for remote access to quantum computers but also serve as the strategic core for achieving large-scale quantum information processing, forming the foundational infrastructure for the future global-scale quantum internet. Quantum teleportation, which enables the transmission of unknown quantum states over long distances by employing quantum entanglement together with classical communication, is essential for the distribution of quantum resources in the construction of the global-scale quantum internet. To realize a global-scale quantum internet, quantum repeater protocols represent one of the most promising approaches for enabling quantum communication between any nodes. This concise review presents representative experimental demonstrations of quantum teleportation for constructing quantum networks across different physical platforms. Along this trajectory, the review discusses current challenges, open issues, and future perspectives toward scalable and practical quantum internet.

quant-ph

Enhancing Facial Expression Recognition in Head-Mounted Displays with Synthetic Data

Facial expression recognition (FER) is crucial for social interaction in mixed reality environments that employ head-mounted displays (HMD). However, collecting FER data from head-mounted cameras (HMC) is challenging due to privacy concerns and the diversity of HMD platforms. Moreover, existing FER datasets are not directly applicable due to the unique perspectives of HMCs. The lack of sufficient data hinders the development of neural network-based HMC FER methods. To address data scarcity, we propose a data synthesis framework that generates HMC-view images from frontal-view images, leveraging abundant existing annotated datasets. Specifically, we first reconstruct 3D textured meshes from images and then apply a configurable camera system to render images from the HMC perspective. Additionally, we introduce a texture-space alignment network (TSAN) that enables accurate texture sampling from images to preserve detailed facial expressions. To evaluate the proposed method, we conduct extensive experiments on both simulated and real HMC datasets. Experimental results demonstrate that models trained on our synthetic dataset outperform those trained on existing datasets and exhibit better generalization across different camera configurations.

cs.CV

Open-Weather Robust 3D Detection via Dual-Critic Diffusion Alignment

Robust 3D object detection under adverse weather remains a critical hurdle for autonomous driving. Despite progress with LiDAR-4D radar fusion, most methods are constrained by a closed-world assumption, implicitly requiring training and test weather to align in both type and severity. This premise fails in practice: the open-ended nature of weather, and even variations within a single type like rain, cause dramatically different LiDAR degradation patterns, leading to significant performance drops in unseen conditions. To address this, we present Dual-Critic Guided Diffusion Alignment (DCDA), a weather-agnostic framework that learns to recover degraded LiDAR features toward a clean manifold. Rather than modeling specific weather types, DCDA employs a 4D radar-conditioned diffusion process to progressively refine features, guided by two complementary critics. (i) A detection-guided critic, anchored by a pre-trained clean-weather model, ensures that the refined features retain object-level discriminability and localization accuracy. (ii) A weather adversarial critic enforces holistic distributional consistency with clean-weather representations. By aligning features through semantic and distributional constraints rather than explicit weather modeling, DCDA generalizes effectively to unseen weather types and severities without requiring paired data or weather labels. We further introduce a structured open-weather benchmark with held-out type-severity combinations and extensive experiments verify DCDA's advantages.

cs.CV

Quantum LiDAR with non-local modulation

Quantum light detection and ranging (LiDAR) utilizes quantum entanglement and correlation to improve precision, noise resilience and covertness of target detection. Despite recent advances, the development of a quantum LiDAR system that simultaneously achieves high precision and a large measurement range remains challenging. Here, we demonstrate a quantum amplitude-modulated continuous wave LiDAR with micrometer precision achievable via increased acquisition time and meter-scale measurement range. In our demonstration, the signal photons directly illuminate the target, while the idler photons are non-locally modulated with a high-frequency cosine wave and never interact with the target. By leveraging the non-local modulation and the quantum correlation, the target detection is achieved with a precision of 0.64 $\pm$ 0.06 mm within one second over a measurement range of 2-8 m. As the acquisition time is up to 500 s, the system achieves a precision of 29 $\pm\ 4{\ \mathrm{\mu m}}$. Furthermore, our system realizes a 50 times precision improvement over the classical single-photon scheme in a background noise 37 dB stronger than the returned probe photons. With these advantages, our method will open venues for the development of high-precision, long-range, and noise-resilient target detection.

quant-ph

Quantum light source with lithium tantalate for scalable photonic quantum circuits

Thin-film lithium tantalate (TFLT) has emerged as a promising integrated photonic platform owing to its low photorefractive noise, high optical damage threshold, and reduced birefringence, attracting increasing interest for scalable photonic technologies. Here, to the best of our knowledge, we demonstrate the first quantum light source with TFLT via spontaneous four-wave mixing, bridging the gap between the rapidly advancing classical TFLT ecosystem and integrated quantum photonics. The fabricated microring exhibits a free spectral range of 350~GHz and an optical quality factor of $10^6$, enabling efficient cavity-enhanced nonlinear interactions. Correlated photon pairs are generated across the telecom band from 1510 to 1570~nm, with a photon pair generation rate of 24 $\mathrm{MHz/mW^{2}}$ at a wavelength of 1535.04 nm. The source delivers strongly antibunched heralded single photons with $g^{(2)}_{H}(0)=0.071\pm0.004$ at a heralding rate of 170 kHz, while the unheralded statistics yield $g^{(2)}(0)=1.93 \pm 0.05$, indicating near-single-temporal-mode emission. Energy-time entanglement is further confirmed by a raw two-photon interference visibility of $92.55\pm0.94\%$, well above the Bell-inequality violation threshold. These results establish TFLT as a manufacturing-compatible platform for scalable photonic quantum circuits, paving the way for the monolithic co-integration of classical and quantum photonic functionalities.

quant-ph

Integrated time-bin entangled quantum light source on a 4H-SiC microring chip

Integrated time-bin-entangled photon-pair source with cavity-enhanced nonlinear optical processes is essential for quantum information technologies. However, microcavities with a high quality factor inherently introduce a trade-off between generation efficiency and photon bandwidth, which hinders the development of high-speed quantum networks with an integrated source. Here, we address this challenge by optimizing the nonlinearity property of the material and the geometry of the integrated microring resonator with a 4H-silicon carbide platform. Operating at a loaded quality factor of 1.9 $\times$ 10^5 - spectral bandwidth of 1.0 GHz and pumped with 300-ps double pulses separated by 1.25 ns at a repetition rate of 160 MHz, the device achieves a time-bin-entangled photon-pair generation rate of 1.35 $\times$ 10^7 s^-1 mW^-2. A raw visibility of 95.55 $\pm$ 0.18% is measured, showing a violation of Bell's inequality by more than 138 standard deviations, and a fidelity of 94.37 $\pm$ 0.22% is obtained by quantum state tomography. These results provide a scalable pathway to an efficient and broadband time-bin entangled quantum light source, overcoming intrinsic limitations of cavity-based designs and advancing integrated platforms for future quantum communication networks.

quant-ph

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance the reasoning-grounded planning for image editing, we propose DDA-Thinker, a Thinker-centric framework designed for the independent optimization of a planning module (Thinker) over a fixed generative model (Editor). This decoupled Thinker-centric paradigm facilitates a controlled analysis of the planning module and makes its contribution under a fixed Editor easier to assess. To effectively guide this Thinker, we introduce a dual-atomic reinforcement learning framework. This framework decomposes feedback into two distinct atomic rewards implemented through verifiable checklists: a cognitive-atomic reward to directly assess the quality of the Thinker's executable plan, which serves as the actionable outcome of the Thinker's reasoning, and a visual-atomic reward to assess the final image quality. To improve checklist quality, our checklist synthesis is grounded not only in the source image and user instruction but also in a rational reference description of the ideal post-edit scene. To support this training, we further develop a two-stage data curation pipeline that first synthesizes a diverse and reasoning-focused dataset, then applies difficulty-aware refinement to curate an effective training curriculum for reinforcement learning. Extensive experiments on reasoning-driven image editing benchmarks, including RISE-Bench and KRIS-Bench, demonstrate that our approach substantially improves overall performance. Our method enables a community model to achieve results competitive with strong proprietary models, highlighting the practical potential of Thinker-centric optimization under a fixed-editor setting.

cs.CV

Change-of-Rings Theorems for the Small Finitistic Dimension

In this paper, we study the small finitistic dimension of a commutative ring from the viewpoint of finitistic flat homological algebra. Using the class $FPR(R)$ of modules admitting finite projective resolutions, we investigate the finitistic flat ($FT$-flat) dimension and establish several of its basic properties. We prove change-of-rings results for the $FT$-flat dimension, including quotient and polynomial extension results, as well as localization inequalities. As applications, we obtain characterizations of the small finitistic dimension in terms of $FT$-flat dimension, derive quotient and polynomial extension theorems for the small finitistic dimension, and establish local upper bounds in terms of the small finitistic dimensions of localizations.

math.AC

CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection

Large language model agents rely on effective model context to obtain task-relevant information for decision-making. Many existing context engineering approaches primarily rely on the context generated from the past experience and retrieval mechanisms that reuse these context. However, retrieved context from past tasks must be adapted by the execution agent to fit new situations, placing additional reasoning burden on the underlying LLM. To address this limitation, we propose a generative context augmentation framework using Contrastive Learning of Experience via Agentic Reflection (CLEAR). CLEAR first employs a reflection agent to perform contrastive analysis over past execution trajectories and summarize useful context for each observed task. These summaries are then used as supervised fine-tuning data to train a context augmentation model (CAM). Then we further optimize CAM using reinforcement learning, where the reward signal is obtained by running the task execution agent. By learning to generate task-specific knowledge rather than retrieve knowledge from the past, CAM produces context that is better tailored to the current task. We conduct comprehensive evaluations on the AppWorld and WebShop benchmarks. Experimental results show that CLEAR consistently outperforms strong baselines. It improves task completion rate from 72.62% to 81.15% on AppWorld test set and averaged reward from 0.68 to 0.74 on a subset of WebShop, compared with baseline agent. Our code is publicly available at https://github.com/awslabs/CLEAR.

cs.AI

SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited attention to fine-grained spatial relationships, often producing images that appear plausible overall yet contain inaccuracies in object positioning. In this work, we present \textbf{SpatialReward}, a verifiable reward model explicitly designed to evaluate spatial layouts in generated images. SpatialReward adopts a multi-stage pipeline: a \emph{Prompt Decomposer} extracts entities, attributes, and spatial metadata from free-form prompts; expert detectors provide accurate visual grounding of object positions and attributes; and a vision-language model applies chain-of-thought reasoning over grounded observations to assess complex spatial relations that are challenging for rule-based methods. To more comprehensively evaluate spatial relationships in generated images, we introduce \textbf{SpatRelBench}, a benchmark covering object attributes, orientation, inter-object relations, and rendered text placement. Experiments on Stable Diffusion and FLUX show that incorporating SpatialReward into RL training consistently improves spatial consistency and overall generation quality, with results aligned more closely to human judgments. These findings indicate that verifiable reward models hold considerable potential for enabling more accurate and controllable optimization in text-to-image generation models.

cs.CV

Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning

Video large language models have demonstrated remarkable capabilities in video understanding tasks. However, the redundancy of video tokens introduces significant computational overhead during inference, limiting their practical deployment. Many compression algorithms are proposed to prioritize retaining features with the highest attention scores to minimize perturbations in attention computations. However, the correlation between attention scores and their actual contribution to correct answers remains ambiguous. To address the above limitation, we propose a novel \textbf{C}ontribution-\textbf{a}ware token \textbf{Co}mpression algorithm for \textbf{VID}eo understanding (\textbf{CaCoVID}) that explicitly optimizes the token selection policy based on the contribution of tokens to correct predictions. First, we introduce a reinforcement learning-based framework that optimizes a policy network to select video token combinations with the greatest contribution to correct predictions. This paradigm shifts the focus from passive token preservation to active discovery of optimal compressed token combinations. Secondly, we propose a combinatorial policy optimization algorithm with online combination space sampling, which dramatically reduces the exploration space for video token combinations and accelerates the convergence speed of policy optimization. Extensive experiments on diverse video understanding benchmarks demonstrate the effectiveness of CaCoVID. Codes are available at https://github.com/LivingFutureLab/CaCoVID.

cs.CV

Unified Thinker: A General Reasoning Modular Core for Image Generation

Despite impressive progress in high-fidelity image synthesis, generative models still struggle with logic-intensive instruction following, exposing a persistent reasoning--execution gap. Meanwhile, closed-source systems (e.g., Nano Banana) have demonstrated strong reasoning-driven image generation, highlighting a substantial gap to current open-source models. We argue that closing this gap requires not merely better visual generators, but executable reasoning: decomposing high-level intents into grounded, verifiable plans that directly steer the generative process. To this end, we propose Unified Thinker, a task-agnostic reasoning architecture for general image generation, designed as a unified planning core that can plug into diverse generators and workflows. Unified Thinker decouples a dedicated Thinker from the image Generator, enabling modular upgrades of reasoning without retraining the entire generative model. We further introduce a two-stage training paradigm: we first build a structured planning interface for the Thinker, then apply reinforcement learning to ground its policy in pixel-level feedback, encouraging plans that optimize visual correctness over textual plausibility. Extensive experiments on text-to-image generation and image editing show that Unified Thinker substantially improves image reasoning and generation quality.

cs.CV