arXiv ScienceSearch

arXiv subjects

Zhisheng Yang

Publications and source records attributed to Zhisheng Yang.

6 recordsLinked to original sources

EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance

Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), has advanced LLM reasoning. However, GRPO suffers from three credit assignment failures: uniform token-level granularity that ignores heterogeneous informational value, uniform polarity that penalizes correct steps and rewards incorrect ones, and zero-variance collapse that erases outcome-driven gradients. We systematically quantify these failures, revealing highly non-uniform token informativeness, widespread step-level polarity misalignment, and substantial training waste. To address these limitations, we propose Entropy-Progress Aligned GRPO (EP-GRPO), a framework that mines the model's intrinsic information flow for dense, self-supervised guidance. EP-GRPO integrates entropy-gated modulation to prioritize high entropy decision pivots, implicit process signals from policy divergence anchored to outcome advantages for directional token-level feedback without external reward models, and cumulative entropy mapping that enables progress-aligned advantage normalization, naturally maintaining gradient flow under zero reward variance. Extensive experiments on mathematical reasoning benchmarks demonstrate that EP-GRPO achieves superior accuracy and efficiency compared to GRPO and its variants. The code will be available.

cs.LG

ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models

Reinforcement learning from verifiable rewards has significantly advanced the reasoning capabilities of large language models. However, Group Relative Policy Optimization (GRPO) typically assigns a uniform, sequence-level advantage to all tokens, thereby overlooking the intrinsic information heterogeneity along reasoning chains. We show that this coarse-grained credit assignment leads to premature entropy collapse and encourages the model to generate redundant, low-quality reasoning paths. Through systematic empirical analysis, we identify Critical Decision Pivots (CDPs): transient high-entropy states where the policy's trajectory is most sensitive to perturbations. These pivots represent the "forks in the road" where effective multi-path exploration is most crucial yet often suppressed by uniform advantage signals. Building on these insights, we propose Entropy-Regulated Policy Optimization (ERPO), which transitions the optimization focus from coarse sequences to fine-grained token dynamics. ERPO introduces three synergistic components: (i) Entropy-aware Gating, which adaptively amplifies exploration at CDPs to facilitate diverse path discovery; (ii) Bucket-based Implicit Normalization, which mitigates difficulty bias by aligning token progress windows; and (iii) Result-anchored Advantage Synthesis, which re-weights token-level signals via outcome-driven anchors. Extensive experiments on competitive mathematical benchmarks demonstrate that ERPO significantly outperforms GRPO. Notably, ERPO not only boosts reasoning accuracy but also yields significantly more concise and robust derivation paths, while achieving performance comparable to large models with orders of magnitude more parameters.

cs.LG

Pulsed heterodyne Brillouin detection enables high-resolution epi-detected biomechanical microscopy and endoscopy

Brillouin microscopy enables non-contact, three-dimensional mapping of viscoelasticity in living systems, yet two long-standing limitations have constrained its biological reach: the lack of high-spectral-resolution epi-detection and the absence of practical fiber-compatible implementations. Here we introduce pulsed heterodyne Brillouin detection (PHBD), a coherent time-domain scheme addressing both challenges. By combining high-peak-power pulsed excitation with shot-noise-limited detection, PHBD reduces the optical dose by approximately two orders of magnitude relative to continuous-wave heterodyne approaches. In an epi-microscope configuration, PHBD attains a spectral resolution of 27 MHz, a tenfold improvement over state-of-the-art Brillouin microscopes, enabling high-specificity, low-phototoxicity imaging of live cells and complex tissues. In an endoscopic configuration, coherent gating rejects parasitic Brillouin background from the delivery fiber, accelerating acquisition by two to three orders of magnitude over previous fiber-optic Brillouin endoscopes. Together, these capabilities establish a unified platform for single-ended, fiber-compatible Brillouin biomechanics, extending mechanical imaging and spectroscopy from cells to deep tissues via minimally invasive probes.

physics.optics

Enhance Large Language Models as Recommendation Systems with Collaborative Filtering

As powerful tools in Natural Language Processing (NLP), Large Language Models (LLMs) have been leveraged for crafting recommendations to achieve precise alignment with user preferences and elevate the quality of the recommendations. The existing approaches implement both non-tuning and tuning strategies. Compared to following the tuning strategy, the approaches following the non-tuning strategy avoid the relatively costly, time-consuming, and expertise-requiring process of further training pre-trained LLMs on task-specific datasets, but they suffer the issue of not having the task-specific business or local enterprise knowledge. To the best of our knowledge, none of the existing approaches following the non-tuning strategy explicitly integrates collaborative filtering, one of the most successful recommendation techniques. This study aims to fill the gap by proposing critique-based LLMs as recommendation systems (Critic-LLM-RS). For our purpose, we train a separate machine-learning model called Critic that implements collaborative filtering for recommendations by learning from the interactions between many users and items. The Critic provides critiques to LLMs to significantly refine the recommendations. Extensive experiments have verified the effectiveness of Critic-LLM-RS on real datasets.

cs.IR

A Framework for Spontaneous Brillouin Noise: Unveiling Fundamental Limits in Brillouin Metrology

Spontaneous Brillouin scattering (SpBS) provides a non-contact tool for probing the mechanical and thermodynamic properties of materials, enabling important applications such as distributed optical fiber sensing and high-resolution Brillouin microscopy. Achieving metrological precision in these systems relies critically on identifying fundamental noise sources. While a pioneering study three decades ago numerically investigated an intrinsic SpBS noise mechanism, this phenomenon has remained largely unexplored, particularly in the context of Brillouin metrological systems. Here, by revisiting its physical formation process and rethinking its stochastic behaviors, we develop and experimentally validate a comprehensive analytical framework on this long-overlooked noise source. Importantly, we theoretically predict, for the first time, the SpBS noise is a universal and fundamental limit that can dominate over conventional limits such as shot noise in Brillouin metrological systems like imaging, microscopy and sensing. Specifically, we experimentally demonstrate the SpBS-noise-limited regime in Brillouin imaging and sensing scenarios. This framework establishes a critical foundation for understanding and optimizing the performance bounds of current and future Brillouin-based technologies across diverse applications.

physics.optics

Integrated photonics modular arithmetic processor

Integrated photonics computing has emerged as a promising approach to overcome the limitations of electronic processors in the post-Moore era, capitalizing on the superiority of photonic systems. However, present integrated photonics computing systems face challenges in achieving high-precision calculations, consequently limiting their potential applications, and their heavy reliance on analog-to-digital (AD) and digital-to-analog (DA) conversion interfaces undermines their performance. Here we propose an innovative photonic computing architecture featuring scalable calculation precision and a novel photonic conversion interface. By leveraging Residue Number System (RNS) theory, the high-precision calculation is decomposed into multiple low-precision modular arithmetic operations executed through optical phase manipulation. Those operations directly interact with the digital system via our proposed optical digital-to-phase converter (ODPC) and phase-to-digital converter (OPDC). Through experimental demonstrations, we showcase a calculation precision of 9 bits and verify the feasibility of the ODPC/OPDC photonic interface. This approach paves the path towards liberating photonic computing from the constraints imposed by limited precision and AD/DA converters.

physics.optics