arXiv ScienceSearch

arXiv subjects

Ruike Cao

Publications and source records attributed to Ruike Cao.

8 recordsLinked to original sources

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.

cs.LG

RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

Long-running LLM agents require memory that persists and evolves across sessions. Text-based memory retrieves and reconstructs past interactions at every query, making long-horizon performance increasingly dependent on retrieval quality and contextual reasoning as histories grow. Parametric memory encodes experience directly into model computation, but existing approaches provide limited support for cross-session memory evolution. Their coupling to a specific backbone further restricts memory reuse after model replacement. We introduce RPMem, a two-stage architecture that compiles each session into a model-independent latent memory through forward computation and selectively integrates it with retained memory via a task-trained recurrent gate. The consolidated memory is then mapped to backbone-specific low-rank adaptation (LoRA) parameters, allowing the encoding capability to transfer when the backbone is replaced. Evaluation across three long-term memory benchmarks and five diverse backbones demonstrates broad generalization with near-constant update cost and memory footprint. With Qwen3-8B on PERMA, RPMem reaches 85.52%, outperforming the strongest parametric and text-based baselines by 5.32 and 12.98 percentage points, respectively. Ablations validate the complementary roles of session compilation and cross-session consolidation, while dynamics analyses reveal that the gate acquires task-specific memory integration strategies. These results establish RPMem as a lifecycle-independent parametric memory framework that maintains evolving cross-session memory that remains reusable across backbone replacements. Our implementation is available at https://github.com/Quark-Medical/rpmem/tree/main.

cs.CL

Design and Evaluation of a PMT High-Voltage system for Deepsea Neutrino Telescope

We present the design and characterization of a Cockcroft--Walton (CW) high-voltage (HV) system developed for deep-sea neutrino telescopes. The system provides independently adjustable bias voltages for 31 three-inch photomultiplier tubes (PMTs) housed in a hybrid Digital Optical Module (hDOM). We describe the system architecture, control logic, and laboratory test procedures, and report the combined PMT--base performance in terms of baseline stability, gain uniformity, and timing accuracy under conditions designed to emulate the deep-sea environment. Baseline measurements show low and stable electronic noise. Gain calibrations based on single-photoelectron spectra demonstrate that all PMTs can be tuned to a common nominal gain and remain stable over multi-day operation. Transit-time-spread measurements yield values below 1.8~ns (FWHM), consistent with manufacturer specifications. These results indicate that the CW-based HV system provides the stability and timing precision required for deep-sea multi-PMT optical modules.

hep-ex

ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical Dialogue

Effective information seeking in multi-turn medical dialogues is critical for accurate diagnosis, especially when dealing with incomplete information. Aligning Large Language Models (LLMs) for these interactive scenarios is challenging due to the uncertainty inherent in user-agent interactions, which we formulate as a Hierarchical Markov Decision Process (H-MDP). While conventional Reinforcement Learning (RL) methods like Group Relative Policy Optimization (GRPO) struggle with long-horizon credit assignment and Proximal Policy Optimization (PPO) suffers from unstable value estimation in this context, we propose a novel uncertainty-aware Adaptive Tree Policy Optimization (ATPO) algorithm. Our method adaptively allocates the rollout budget to states with high uncertainty, quantified by a composite metric of Bellman error and action-value variance. This strategy enables more accurate value estimation, while fostering more efficient and diverse exploration. To mitigate the high computational cost of tree-based RL, we introduce two key optimizations: an uncertainty-guided pruning mechanism to minimize the number of rollouts, and an asynchronous search architecture that leverages KV cache reuse to maximize inference throughput. Extensive experiments on three public medical dialogue benchmarks demonstrate that our algorithm significantly outperforms several strong baselines, culminating in Qwen3-8B model surpassing the much larger GPT-4o ($+0.92\%$ accuracy).

cs.LG

A Cost Effective Optimization of the hybrid-DOM Design for TRIDENT

TRIDENT is a planned multi-cubic-kilometer deep-sea neutrino telescope to be built in the South China Sea, designed to rapidly discover high-energy astrophysical neutrino sources with sensitivity to all neutrino flavors. Achieving this at scale requires a detector design that balances performance with power, cost, and mechanical simplicity. This study presents a cost-effective optimization of TRIDENT's hybrid Digital Optical Module (hDOM) design, comparing configurations using high-quantum-efficiency (QE) 3-inch PMTs and larger 4-inch PMTs, the latter evaluated with both baseline and enhanced QE assumptions. Using full-chain detector simulations incorporating site-specific seawater optical properties and realistic backgrounds, we assess performance in all-flavor neutrino detection efficiency, directional reconstruction, and tau neutrino flavor identification from 1 TeV to 10 PeV. We find that if 4-inch PMTs can achieve QE comparable to 3-inch PMTs, their performance matches or improves upon that of the 3-inch design, while significantly reducing channel count, power consumption, and cost. These findings support the 4-inch PMT hDOM as a promising and scalable choice for TRIDENT's future instrumentation.

hep-ex

MuonSLab: A plastic scintillator based detector for muon measurement in the deep ocean

Atmospheric muons are important probes for studying primary cosmic rays and extensive air showers. Additionally, they constitute a significant background for many underground and deep-sea neutrino experiments, such as TRopIcal DEep-sea Neutrino Telescope (TRIDENT). Understanding the muon flux at various depths in the deep sea is essential for validating TRIDENT simulations and guiding the development of optimized trigger strategies. This paper introduces a novel device based on plastic scintillalors and silicon photomultipliers (SiPMs) named MuonSLab, which is designed to measure muon flux in the deep sea and has the potential to be extended to other atmospheric muon property measurements. We discuss the design and instrumentation of MuonSLab and present results from several muon flux measurements, demonstrating its sensitivity to muon detection and its stability during operations across multiple locations.

hep-ex

Front-end electronics development of large-area SiPM arrays for high-precision single-photon time measurement

TRopIcal DEep-sea Neutrino Telescope (TRIDENT) plans to incorporate silicon photomultipliers (SiPMs) with superior time resolution in addition to photomultiplier tubes (PMTs) into its detection units, namely hybrid Digital Optical Modules (hDOMs), to improve its angular resolution. However, the time resolution significantly degrades for large-area SiPMs due to the large detector capacitance, posing significant challenges for the readout electronics of SiPMs in hDOM. We analyzed the influences of series and parallel connections when constructing a large-area SiPM array and designed a series-parallel connection SiPM array with differential output. We also designed a high-speed pre-amplifier based on transformers (MABA-007159) and radio frequency amplifiers (BGA2803), and an analog multi-channel summing circuit based on operational amplifiers (LMH6629). We measured the single photon time resolution (SPTR) of a $4\times4$ SiPM (Hamamatsu S13360-3050PE) array ($12\times12~\mathrm{mm}^2$) of approximately 300 ps FWHM. This front-end readout design enables the large-area SiPM array to achieve high-precision single photon time measurement in one readout channel.

physics.ins-det

Nonabelian Kinetic Mixing in a Confining Phase

Dark matter from a hidden sector with SU($N$) gauge symmetry can have a nonabelian kinetic mixing portal with the standard model. The dark photon becomes massive in the confining phase without the need for spontaneous symmetry breaking. Depending on the particle content of the dark sector, there can be two or more composite vectors that get kinetic mixing through a heavy mediator particle $X$. This provides a model of composite dark photons giving a portal for direct detection of dark baryons. Avoiding exotic charged relics requires additional couplings allowing $X$ to decay to dark quarks and standard model fields, leading to further portals between the dark matter and the standard model. We comprehensively study the constraints on such models from colliders, rare decays, direct detection, and big bang nucleosynthesis.

hep-ph