arXiv ScienceSearch

arXiv subjects

Zhenrong Zhang

Publications and source records attributed to Zhenrong Zhang.

At least 19 recordsLinked to original sources

ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning

Reinforcement learning has emerged as an effective paradigm for enhancing the mathematical reasoning capabilities of large language models. Among existing policy optimization methods, Proximal Policy Optimization (PPO) remains particularly appealing because its learned critic can, in principle, provide token-level credit assignment. However, in mathematical reasoning tasks characterized by long reasoning horizons and sparse outcome rewards, reliable token-level credit assignment remains challenging. The standard critic often fails to accurately evaluate intermediate reasoning states, resulting in noisy advantage estimates and suboptimal policy updates. In this paper, we propose ReDiPPO, a Reference-guided and Discrepancy-aware PPO framework for mathematical reasoning. ReDiPPO introduces a reference-guided critic that uses reference answers as training-time privileged signals to provide more accurate value estimation. Meanwhile, it retains a standard critic and quantifies the token-level reference-standard discrepancy between the standard value estimate and the reference-guided value estimate. This discrepancy serves as an indicator of difficult reasoning states and is used to reweight the corresponding token-level advantages during PPO optimization. Extensive experiments on diverse mathematical reasoning benchmarks demonstrate that ReDiPPO improves value-estimation accuracy and consistently outperforms strong policy optimization baselines, including PPO, DAPO, and GSPO, in final reasoning performance. Our code is available on https://github.com/cii030/ReDiPPO.

cs.AI

A fault-tolerant quantum blockchain deployed on commercial telecommunications network

Popularized by the Bitcoin cryptocurrency, blockchain technology establishes a decentralized digital framework that utilizes cryptographic and consensus protocols to secure data against unauthorized modification. Consequently, blockchain has found broad adoption across diverse fields, including finance, data management, healthcare, and digital asset governance. In the quantum computing era, a paramount objective for blockchain is to preserve its foundational advantages of cryptographic integrity and decentralized fault-tolerant resilience. In principle, quantum digital signatures and quantum Byzantine agreement protocols offer foundational security guarantees and tolerate up to one-half of malicious nodes for blockchain. However, the practical realization of such a quantum-enhanced blockchain remains a significant and multifaceted challenge. Here, we propose and experimentally demonstrate a fully operational hybrid quantum blockchain architecture built on photonic integrated circuits and deployed over commercially available classical telecommunications infrastructure. The system achieves a fault tolerance of nearly one-half, surpassing the classical limit, while reaching consensus on a timescale of seconds. A deployed food traceability application validates the practicality of the proposed architecture, achieving a throughput of approximately 500 transactions per second. This work establishes a foundation for practical quantum blockchains, enabling secure, scalable, and decentralized information processing in the emerging quantum era.

quant-ph

Experimental demonstration of scalable quantum blockchain with exponentially superior quantum communication complexity

To secure modern distributed digital infrastructures, quantum blockchains exploit quantum resources to achieve information-theoretic security and surpass the classical one-third fault-tolerance bound. However, existing high-fault-tolerant protocols face a fundamental scalability challenge: the blockchain trilemma imposes either exponential communication complexity or experimentally demanding multipartite entanglement. Here, we experimentally demonstrate a scalable quantum blockchain protocol based on weak coherent states that achieves an exponential reduction in quantum communication complexity. The protocol employs a circular quantum Byzantine agreement mechanism that preserves information-theoretic security while avoiding multipartite entanglement. We implement this protocol on a photonic integrated circuit platform, realizing a six-node network over commercially available telecommunication infrastructure. Compared with previous schemes, the protocol requires less than 4% of the quantum communication resources. Leveraging this advantage, we further demonstrate a quantum-secured token exchange application achieving a throughput of 805.3 transactions per second with zero failures. These results establish a practical pathway toward scalable quantum blockchain.

quant-ph

Inverse-squeezing receivers for squeezed-state pulse-position modulation under ideal and phase-diffusion conditions

We introduce a squeezed-state pulse-position modulation (S-PPM) format, where the empty slots are squeezed vacuum states and the pulse slot is a displaced squeezed state. Based on this property, we propose an inverse-squeezing conditional pulse-nulling (IS-CPN) receiver. In the ideal case, inverse squeezing maps S-PPM into an equivalent coherent-state PPM signal with a large pulse energy, leading to a closed-form expression for the receiver error probability. We further analyze IS-CPN under common phase diffusion using a finite-path MAP formulation with phase-averaged likelihoods. Numerical results show that IS-CPN outperforms conventional CPN under the same energy constraint and remains advantageous under phase noise and finite photon-number resolution. These results demonstrate that combining squeezed-state modulation with inverse-squeezing conditional nulling can improve photon-efficient optical communication.

quant-ph

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

Large Language Models (LLMs) show strong potential as educational tutors. Existing approaches typically train them to solve problems and provide correct answers, but this problem-solving-centered paradigm overlooks key requirements of effective tutoring: progressive guidance and the coordination of multiple pedagogical objectives across multi-turn interactions. Developing such tutors remains challenging because student behavior varies substantially with individual knowledge states, pedagogical effectiveness depends on multiple factors beyond final-answer correctness, and coordinating these objectives over tutor-student interactions is inherently difficult. To address these challenges, we propose PEARL, a PEdagogically Aligned Reinforcement Learning framework for training Socratic tutoring agents. First, we introduce a controllable student simulator that disentangles latent cognitive states from response generation, enabling simulation of diverse abilities and misconceptions. Second, we develop a pedagogically aligned reward model that jointly assesses pedagogical quality and objective correctness. Finally, we propose a stable multi-objective reinforcement learning approach that balances competing pedagogical objectives during tutor training. Experiments across multiple benchmarks show that PEARL performs competitively against tutoring-specific open-source systems and leading proprietary LLMs.

cs.LG

Residual-Squeezing Mechanism of Mismatch in Inverse-Squeezing Kennedy Receivers

The discrimination of quantum states is fundamental to quantum information processing. Inverse-squeezing Kennedy (IS-Kennedy) receivers can outperform the coherent-state BPSK Helstrom benchmark at the same energy by converting transmitter-side squeezing into an effective coherent-state separation gain, without violating the Helstrom bound for the squeezed-state alphabet. This work investigates how squeezing mismatch degrades this mechanism. We show that imperfect inverse squeezing transforms the ideally nulled output into a residually squeezed state, thereby altering the photon-number statistics before detection. This residual-squeezing picture reveals a strong physical asymmetry between squeezing-magnitude and squeezing-phase mismatches. Magnitude mismatch produces an energy-independent error floor in the high-signal-energy regime, whereas phase mismatch generates a residual squeezing term that grows with signal energy. In the small-residual-squeezing regime, this leads to a polynomial growth of the leading error contribution and a rapid collapse of the SQL advantage. We also identify a parity-step effect in photon-number-resolving detection: because the nulled residual squeezed vacuum contains only even photon numbers, increasing detector resolution improves the high-energy robustness only when the effective saturation threshold crosses the next even photon number. These results identify phase locking as the dominant bottleneck for IS-Kennedy-type non-Gaussian receivers under unitary squeezing mismatch and provide design guidelines for robust squeezed-state quantum receivers.

quant-ph

Near-optimal discrimination of displaced squeezed binary signals using displacement, inverse-squeezing, and photon-number-resolving detection

We propose an inverse-squeezing Kennedy receiver for discriminating binary phase-shift-keyed displaced squeezed vacuum states. The receiver combines a Kennedy-type nulling displacement, an orthogonally oriented inverse-squeezing operation and photon-number-resolving detection with a maximum-a-\emph{posteriori} threshold rule. Its key mechanism is that the inverse-squeezing stage converts transmitter-side squeezing into enhanced photon-number contrast, or equivalently an effective coherent-state energy gain, that can be directly exploited at the measurement stage. Under ideal equal-prior conditions, the receiver surpasses the standard quantum limit for squeezed-state binary phase-shift keying at approximately $N\approx 0.3$, outperforms the Helstrom bound of coherent-state binary phase-shift keying at approximately $N\approx 0.4$, and reaches the 1\% error level near $N\approx 0.6$. We further analyze its performance under realistic imperfections, including finite detector efficiency, dark counts, channel phase diffusion, receiver thermal noise and transmission loss. The results show that adaptive thresholding preserves robust performance against detector and noise imperfections over practical parameter ranges, whereas transmission loss progressively suppresses the squeezing-enabled advantage. These findings indicate that, for the fixed source parametrization adopted in this work, the proposed receiver is most advantageous in the low-loss regime, especially at low source energies.

quant-ph

High-Rate Free-Running Reference-Frame-Independent Measurement-Device-Independent Quantum Key Distribution with Classified Distillation

Reference-frame-independent measurement-device-independent quantum key distribution (RFI-MDI-QKD) eliminates detector side-channel attacks and avoids reference-frame calibration. While its feasibility has been widely demonstrated, existing implementations typically assume fixed or slowly drifting reference-frame misalignment, conditions rarely satisfied outside the laboratory. In realistic environments, rapid and free-running reference-frame variations can severely degrade both the key rate and transmission distance of conventional RFI-MDI-QKD. Here we propose a free-running RFI-MDI-QKD protocol that maintains high-rate key generation under rapid reference-frame variations. By introducing a classification-distillation method that reclassifies total detection events, secure keys can be extracted without modifying the experimental setup. Our protocol achieves a key rate more than nine times higher than the best previous RFI-MDI-QKD scheme and tolerates channel losses exceeding 24 dB, where earlier approaches fail. These results enable practical quantum key distribution on mobile platforms, including satellite-to-ground links and airborne nodes.

quant-ph

Step Potential Advantage Estimation: Harnessing Intermediate Confidence and Correctness for Efficient Mathematical Reasoning

Reinforcement Learning with Verifiable Rewards (RLVR) elicits long chain-of-thought reasoning in large language models (LLMs), but outcome-based rewards lead to coarse-grained advantage estimation. While existing approaches improve RLVR via token-level entropy or sequence-level length control, they lack a semantically grounded, step-level measure of reasoning progress. As a result, LLMs fail to distinguish necessary deduction from redundant verification: they may continue checking after reaching a correct solution and, in extreme cases, overturn a correct trajectory into an incorrect final answer. To remedy the lack of process supervision, we introduce a training-free probing mechanism that extracts intermediate confidence and correctness and combines them into a Step Potential signal that explicitly estimates the reasoning state at each step. Building on this signal, we propose Step Potential Advantage Estimation (SPAE), a fine-grained credit assignment method that amplifies potential gains, penalizes potential drops, and applies penalty after potential saturates to encourage timely termination. Experiments across multiple benchmarks show SPAE consistently improves accuracy while substantially reducing response length, outperforming strong RL baselines and recent efficient reasoning and token-level advantage estimation methods. The code is available at https://github.com/cii030/SPAE-RL.

cs.CL

THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning

Large Language Models (LLMs) have made remarkable progress in mathematical reasoning, but still continue to struggle with high-precision tasks like numerical computation and formal symbolic manipulation. Integrating external tools has emerged as a promising approach to bridge this gap. Despite recent advances, existing methods struggle with three key challenges: constructing tool-integrated reasoning data, performing fine-grained optimization, and enhancing inference. To overcome these limitations, we propose THOR (Tool-Integrated Hierarchical Optimization via RL). First, we introduce TIRGen, a multi-agent based pipeline for constructing high-quality datasets of tool-integrated reasoning paths, aligning with the policy and generalizing well across diverse models. Second, to perform fine-grained hierarchical optimization, we introduce an RL strategy that jointly optimizes for both episode-level problem solving and step-level code generation. This is motivated by our key insight that the success of an intermediate tool call is a strong predictor of the final answer's correctness. Finally, THOR incorporates a self-correction mechanism that leverages immediate tool feedback to dynamically revise erroneous reasoning paths during inference. Our approach demonstrates strong generalization across diverse models, performing effectively in both reasoning and non-reasoning models. It further achieves state-of-the-art performance for models of a similar scale on multiple mathematical benchmarks, while also delivering consistent improvements on code benchmarks. Our code will be publicly available at https://github.com/JingMog/THOR.

cs.AI

Scalable twin-field quantum key distribution network enabled by adaptable architecture

Quantum key distribution (QKD) is a key application in quantum communication, enabling secure key exchange between parties using quantum states. Twin-field (TF) QKD offers a promising solution that surpasses the repeaterless limits, and its measurement-device-independent nature makes it suitable for star-type network architectures. In this work, we propose a scalable TF-QKD network with adaptable architecture, where users prepare quantum signals and send them to network nodes. These nodes use an optical switch to route the signals to multi-user measurement units, enabling secure key distribution among arbitrary users and adapting to complex connection demands of the network. A proof-of-principle demonstration with three users successfully achieved secure key sharing over simulated link losses of up to $30$ dB, with an average rate of $19.57$ bit/s. Additionally, simulations show that the proposed architecture can achieve a total secure key rate of $4.84 \times 10^{4}$ bit/s at $100$ km in a symmetric $32$-user network. This approach represents a significant advancement in the topology of untrusted-node QKD networks and holds promise for practical, large-scale applications in secure communication.

quant-ph

Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration

Recent advances in Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general domains and demonstrated promise in multimodal mathematical reasoning. However, applying MLLMs to geometry problem solving (GPS) remains challenging due to lack of accurate step-by-step solution data and severe hallucinations during reasoning. In this paper, we propose GeoGen, a pipeline that can automatically generates step-wise reasoning paths for geometry diagrams. By leveraging the precise symbolic reasoning, \textbf{GeoGen} produces large-scale, high-quality question-answer pairs. To further enhance the logical reasoning ability of MLLMs, we train \textbf{GeoLogic}, a Large Language Model (LLM) using synthetic data generated by GeoGen. Serving as a bridge between natural language and symbolic systems, GeoLogic enables symbolic tools to help verifying MLLM outputs, making the reasoning process more rigorous and alleviating hallucinations. Experimental results show that our approach consistently improves the performance of MLLMs, achieving remarkable results on benchmarks for geometric reasoning tasks. This improvement stems from our integration of the strengths of LLMs and symbolic systems, which enables a more reliable and interpretable approach for the GPS task. Codes are available at https://github.com/ycpNotFound/GeoGen.

cs.CL

MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique

Visual language models (VLMs) have demonstrated strong performance across diverse multimodal reasoning tasks but still face challenges such as hallucinations, resulting in incorrect reasoning outcomes. Inspired by recent research on external feedback mechanisms in large language models (LLMs), we propose a multimodal actor-critic framework to enhance VLM reasoning capabilities. Specifically, the actor model generates step-by-step reasoning paths based on image and text inputs, while the critic model evaluates these reasoning paths and provides corrective feedback. The actor model iteratively refines its reasoning based on the feedback until the reasoning outcome is deemed satisfactory by the critic model. To reduce reliance on costly manual annotations, we introduce an automated method for constructing multimodal critique datasets. By leveraging Monte Carlo Tree Search (MCTS), we systematically guide the actor model to explore diverse reasoning paths. To obtain critique data for correcting erroneous reasoning steps, we prompt an annotator model to compare pairs of reasoning paths diverging from a shared ancestor node - one leading to a correct conclusion and the other to an incorrect one. This approach enables us to construct the MMC (MCTS-based Multimodal Critique) dataset, upon which we further develop a comprehensive training and inference pipeline. Extensive experiments conducted on several public benchmark datasets and mainstream VLMs demonstrate that our approach significantly improves the performance of VLM on complex multimodal reasoning tasks, underscoring its effectiveness and wide applicability.

cs.MM

PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search

Recent work increasingly focuses on improving the reasoning capabilities of Multimodal Large Language Models (MLLMs). Among existing methods, Process Reward Models (PRMs) stand out for offering dense, step-wise supervision to guide intermediate reasoning. However, how to effectively integrate PRMs into search strategies remains an open question. In this paper, we introduce PRM-BAS (PRM-Guided Beam Annealing Search), a lightweight approach for PRM-guided reasoning that dynamically adjusts beam size -- starting with a broader search space and gradually narrowing it as contextual information accumulates, thereby balancing performance and efficiency. We further propose a unified framework for data construction and PRM training. Specifically, we construct the PRM-BAS-300k dataset by selecting 300k questions from existing datasets and performing rollouts at each step to estimate the probability of reaching a correct final answer. The PRM is then trained using a combination of value loss for absolute action quality and rank loss for relative action quality. Extensive experiments on challenging multimodal reasoning benchmarks demonstrate that PRM-BAS significantly improves reasoning performance while maintaining low computational cost. Moreover, it generalizes well across different model scales and architectures, showcasing strong robustness and plug-and-play capability.

cs.MM

Quantum-Secured DSP-Lite Data Transmission Architectures for AI-Driven Data Centres

Artificial intelligence-driven (AI-driven) data centres, which require high-performance, scalable, energy-efficient, and secure infrastructure, have led to unprecedented data traffic demands. These demands involve low latency, high bandwidth connections, low power consumption, and data confidentiality. However, conventional optical interconnect solutions, such as intensity-modulated direct detection and traditional coherent systems, cannot address these requirements simultaneously. In particular, conventional encryption protocols that rely on complex algorithms are increasingly vulnerable to the rapid advancement of quantum computing. Here, we propose and demonstrate a quantum-secured digital signal processing-lite (DSP-Lite) data transmission architecture that meets all the stringent requirements for AI-driven data centre optical interconnects (AI-DCIs) scenarios. By integrating a self-homodyne coherent (SHC) system and quantum key distribution (QKD) through the multicore-fibre-based space division multiplexing (SDM) technology, our scheme enables secure, high-capacity, and energy-efficient data transmission while ensuring resilience against quantum computing threats. In our demonstration, we achieved an expandable transmission capacity of 2 Tbit per second (Tb/s) and a quantum secret key rate (SKR) of 229.2 kb/s, with a quantum bit error rate (QBER) of approximately 1.27% and with ultralow power consumption. Our work paves the way for constructing secure, scalable, and cost-efficient data transmission frameworks, thus enabling the next generation of intelligent, leak-proof optical interconnects for data centres.

quant-ph

Skeleton and Font Generation Network for Zero-shot Chinese Character Generation

Automatic font generation remains a challenging research issue, primarily due to the vast number of Chinese characters, each with unique and intricate structures. Our investigation of previous studies reveals inherent bias capable of causing structural changes in characters. Specifically, when generating a Chinese character similar to, but different from, those in the training samples, the bias is prone to either correcting or ignoring these subtle variations. To address this concern, we propose a novel Skeleton and Font Generation Network (SFGN) to achieve a more robust Chinese character font generation. Our approach includes a skeleton builder and font generator. The skeleton builder synthesizes content features using low-resource text input, enabling our technique to realize font generation independently of content image inputs. Unlike previous font generation methods that treat font style as a global embedding, we introduce a font generator to align content and style features on the radical level, which is a brand-new perspective for font generation. Except for common characters, we also conduct experiments on misspelled characters, a substantial portion of which slightly differs from the common ones. Our approach visually demonstrates the efficacy of generated images and outperforms current state-of-the-art font generation methods. Moreover, we believe that misspelled character generation have significant pedagogical implications and verify such supposition through experiments. We used generated misspelled characters as data augmentation in Chinese character error correction tasks, simulating the scenario where students learn handwritten Chinese characters with the help of misspelled characters. The significantly improved performance of error correction tasks demonstrates the effectiveness of our proposed approach and the value of misspelled character generation.

cs.CV

Experimental secure entanglement-free quantum remote sensing over 50 km of optical fiber

Secure quantum remote sensing (SQRS) uses quantum states to gather information about distant objects or environments while ensuring secure data transmission against eavesdropping. It has potential applications in various fields, including environmental monitoring, military surveillance, and disaster response, where both data accuracy and transmission security are critical. Recent experiments have demonstrated the feasibility of SQRS using entanglement states. Here, we experimentally demonstrate an SQRS that can estimate a phase without requiring entanglement, offering the practical advantage that single-qubit states are easier to prepare. We successfully estimate the preset phase information at a remote site over a fiber distance of 50 km, which serves as a key step toward long-distance applications.

quant-ph

RFL: Simplifying Chemical Structure Recognition with Ring-Free Language

The primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional structures of molecules, particularly those with rings and multiple branches, present significant challenges for current end-to-end methods to learn one-dimensional markup directly. To overcome this limitation, we propose a novel Ring-Free Language (RFL), which utilizes a divide-and-conquer strategy to describe chemical structures in a hierarchical form. RFL allows complex molecular structures to be decomposed into multiple parts, ensuring both uniqueness and conciseness while enhancing readability. This approach significantly reduces the learning difficulty for recognition models. Leveraging RFL, we propose a universal Molecular Skeleton Decoder (MSD), which comprises a skeleton generation module that progressively predicts the molecular skeleton and individual rings, along with a branch classification module for predicting branch information. Experimental results demonstrate that the proposed RFL and MSD can be applied to various mainstream methods, achieving superior performance compared to state-of-the-art approaches in both printed and handwritten scenarios. The code is available at https://github.com/JingMog/RFL-MSD.

cs.CV