arXiv ScienceSearch

arXiv subjects

Yuan He

Publications and source records attributed to Yuan He.

At least 19 recordsLinked to original sources

The classification of some polynomial maps in dimension three

In the paper, we classify all polynomial maps of the form $H=(u(x,y),\allowbreak v(x,y,z),h(x,y,z))$ in the case that $JH$ is nilpotent and $°_zv\geq 2°_zh$. Then we give the structure of $H=(u(x,y),v(x,y,z),h(x,y,z))$ if $JH$ is nilpotent and $°_zh\leq 3$.

math.AC

PIMID: A Full-System Simulator with Intricacy and Diversity for Processing-in-Memory

Processing-in-Memory addresses the memory wall by co-locating computation with memory, but because real PIM hardware remains scarce, simulation is the primary way to explore the PIM design space. Yet existing PIM simulators each cover only part of that space: they typically model a single memory technology, fix processing elements at one level of the memory hierarchy, support a single execution model, and stop at the device boundary. We therefore present PIMID, an execution- and trace-driven full-system simulator that closes these gaps in one tool. PIMID supports both the shared-memory and message-passing execution models, running annotated parallel code in OpenMP and MPI side by side across eleven memory technologies (seven DRAM standards, SRAM, and three non-volatile memories); it places PEs anywhere from subarrays to logic dies, sweeps PE count and core-model fidelity, and prices the in-memory network per technology from measured congestion. Its single-process host-device co-simulation resolves an end-to-end time and energy breakdown (host preparation, device compute, and explicit boundary charges) that device-only tools cannot produce. Across the resulting dual-execution-model dataset, PIMID shows that the memory technology alone moves execution time by more than an order of magnitude and that the best host main memory is not the best PIM substrate; that regular kernels scale superlinearly with PE count as in-memory bandwidth co-scales with compute; that graph traversal under message-passing hits a collective-communication wall absent under shared memory; and that at full-system scope the offload trades time for energy only on the bandwidth-class memory: shared-memory offload saves energy on HBM3 while a 16-core host keeps every end-to-end time win. PIMID's plugin interfaces let new engines and models be added through standardized YAML specifications as PIM technology evolves.

cs.AR

Reciprocity formulas for certain generalized Hardy-Berndt sums

In this paper, we introduce the generalized Hardy-Berndt sums. We establish some formulas of products of the Bernoulli and Euler functions by using the Fourier series technique and some properties of the periodic-zeta and Lerch-zeta functions. As applications of the results, we give some reciprocity formulas for the generalized Hardy-Berndt sums. It turns out that Hardy's reciprocity formulas follow as special cases and that Goldberg's three-term reciprocity formulas are recovered.

math.NT

A unification of Euler-Maclaurin and Euler-Boole summation formulas

In this paper, we establish a summation formula associated with the generalized Apostol-Bernoulli functions. This formula unifies the Euler-Maclaurin summation formula, the Euler-Boole summation formula, and the character analogues of these two formulas. We also apply it to obtain the general power sum formula and the special values of Berndt's generalized $L$-function at non-positive integers.

math.NT

On the special values of Berndt's generalized $L$-function

In this paper, we study the generalized $L$-function considered by Berndt (1975). We introduce the generalized Apostol-Bernoulli polynomials and the generalized Apostol-Bernoulli functions, and establish some properties for them, including the Fourier series for these functions. We show that the values of Berndt's generalized $L$-function at integers are explicitly evaluated in terms of the generalized Apostol-Bernoulli functions.

math.NT

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This paper studies tool-calling along two complementary axes: effectiveness, i.e., how this capability is measured, and efficiency, i.e., how it is learned. On effectiveness, we systematically analyze tool-calling evaluation pipelines and show that results can be highly sensitive to seemingly minor, often undocumented implementation choices including the random seed, system prompt, multi-turn template construction, and how prior interaction/reasoning history is carried forward. These choices can lead to substantial differences in reported performance, especially in multi-turn settings where without rigorous standardization, leaderboard rankings are unreliable. On efficiency, we examine standard reinforcement learning (RL) for tool-calling and identify two sources of computational waste: (i) during rollouts, many prompts produce no learning signal, and (ii) during policy updates, optimization incurs high computational cost. Guided by these findings, we introduce two techniques that accelerate RL-based tool-calling training, achieving substantial wall-clock speedup without degrading performance.

cs.LG

Time-marching representation based quantum algorithms for the Lattice Boltzmann model of the advection-diffusion equation

This article introduces a novel framework for developing quantum algorithms for the Lattice Boltzmann Method (LBM) applied to the advection-diffusion equation. We formulate the collision-streaming evolution of the LBM as a compact time-marching scheme and rigorously establish its stability under low Mach number conditions. This unified formulation eliminates the need for classical measurement at each time step, enabling a systematic and fully quantum implementation. Building upon this representation, we investigate two distinct quantum algorithmic approaches. The first is a time-marching quantum algorithm realized through sequential evolution operators, for which we provide a detailed implementation-including block-encoding and dilating unitarization-along with a full complexity analysis. The second employs a quantum linear systems algorithm, which encodes the entire time evolution into a single global linear system. We demonstrate that both methods achieve comparable asymptotic time complexities. The proposed algorithms are validated through numerical simulations of benchmark problems in one and two dimensions. This work provides a systematic pathway that avoids full-state measurement and reinitialization at every time step for the quantum simulation of advection-diffusion processes via the lattice Boltzmann paradigm.

math-ph

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers

Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforcement Learning with Verifiable Reward (RLVR), particularly in domains like mathematics and programming, where ground-truth correctness can be automatically evaluated. However, extending this success to other reasoning-intensive domains remains challenging due to the scarcity of high-quality, verifiable datasets and the high cost of human supervision. In this work, we introduce the Loong Project: an open-source framework for scalable synthetic data generation and verification across a diverse range of reasoning-intensive domains. The framework consists of two key components: (1) LoongBench, a curated seed dataset containing 8,729 human-vetted examples across 12 domains (e.g., Advanced Mathematics, Chemistry, Logic), each paired with executable code and rich metadata; and (2) LoongEnv, a modular synthetic data generation environment that supports multiple prompting strategies to produce new question-answer-code triples. Together, these components form an agent-environment loop that enables reinforcement learning, where an LLM-based agent is rewarded for generating Chain-of-Thought (CoT) solutions that align with code-executed answers. Empirically, we benchmark LoongBench on a broad suite of both open-source and proprietary LLMs to evaluate domain coverage and reveal performance bottlenecks. In addition, we conduct a comprehensive analysis of synthetic data generated by LoongEnv, examining correctness, difficulty, and diversity. Code and documentation are available at https://github.com/camel-ai/loong.

cs.LG

Spontaneous patterning of cell size on curved surfaces

Tissue surfaces exhibit complex curvature during embryogenesis and oncogenesis. Evidence shows that cells can actively sense curvature to regulate behavior and fate, yet the underlying mechanism remains unclear. Here, we develop a vertex model for arbitrary curved surfaces and uncover spontaneous cell size patterning on ellipsoidal surfaces: cells in high-curvature regions are consistently larger than those in low-curvature regions. This non-uniformity arises from a mechanical competition encoded in Riemannian geometry: positive Gaussian curvature reduces the perimeter-to-area ratio of polygonal cells, relaxing cell-edge tension in high-curvature regions, which is compensated by area expansion to maintain global force balance. This area pattern is robust against variations in model parameters and matches observations in biological systems. The perimeter pattern, in contrast, is governed by competition between the intrinsic geometric tendency and the deformation required by force balance, and undergoes reversal beyond a critical shape index. Together, these findings establish self-organized spatial variations in cell size as a potential physical mechanism for curvature sensing.

physics.bio-ph

Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving

Vision-Language Models (VLMs) offer a promising approach to end-to-end autonomous driving due to their human-like reasoning capabilities. However, troublesome gaps remains between current VLMs and real-world autonomous driving applications. One major limitation is that existing datasets with loosely formatted language descriptions are not machine-friendly and may introduce redundancy. Additionally, high computational cost and massive scale of VLMs hinder the inference speed and real-world deployment. To bridge the gap, this paper introduces a structured and concise benchmark dataset, NuScenes-S, which is derived from the NuScenes dataset and contains machine-friendly structured representations. Moreover, we present FastDrive, a compact VLM baseline with 0.9B parameters. In contrast to existing VLMs with over 7B parameters and unstructured language processing(e.g., LLaVA-1.5), FastDrive understands structured and concise descriptions and generates machine-friendly driving decisions with high efficiency. Extensive experiments show that FastDrive achieves competitive performance on structured dataset, with approximately 20% accuracy improvement on decision-making tasks, while surpassing massive parameter baseline in inference speed with over 10x speedup. Additionally, ablation studies further focus on the impact of scene annotations (e.g., weather, time of day) on decision-making tasks, demonstrating their importance on decision-making tasks in autonomous driving.

cs.CV

EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools

Deep research requires reasoning over web evidence to answer open-ended questions, and it is a core capability for AI agents. Yet many deep research agents still rely on implicit, unstructured search behavior that causes redundant exploration and brittle evidence aggregation. Motivated by Anthropic's "think" tool paradigm and insights from the information-retrieval literature, we introduce Q+, a set of query and evidence processing tools that make web search more deliberate by guiding query planning, monitoring search progress, and extracting evidence from long web snapshots. We integrate Q+ into the browser sub-agent of Eigent, an open-source, production-ready multi-agent workforce for computer use, yielding EigentSearch-Q+. Across four benchmarks (SimpleQA-Verified, FRAMES, WebWalkerQA, and XBench DeepSearch), Q+ improves Eigent's browser agent benchmark-size-weighted average accuracy by 3.0, 3.8, and 0.6 percentage points (pp) for GPT-4.1, GPT-5.1, and Minimax M2.5 model backends, respectively. Case studies further suggest that EigentSearch-Q+ produces more coherent tool-calling trajectories by making search progress and evidence handling explicit.

cs.AI

MetaDAT: Generalizable Trajectory Prediction via Meta Pre-training and Data-Adaptive Test-Time Updating

Existing trajectory prediction methods exhibit significant performance degradation under distribution shifts during test time. Although test-time training techniques have been explored to enable adaptation, current approaches rely on an offline pre-trained predictor that lacks online learning flexibility. Moreover, they depend on fixed online model updating rules that do not accommodate the specific characteristics of test data. To address these limitations, we first propose a meta-learning framework to directly optimize the predictor for fast and accurate online adaptation, which performs bi-level optimization on the performance of simulated test-time adaptation tasks during pre-training. Furthermore, at test time, we introduce a data-adaptive model updating mechanism that dynamically adjusts the predefined learning rates and updating frequencies based on online partial derivatives and hard sample selection. This mechanism enables the online learning rate to suit the test data, and focuses on informative hard samples to enhance efficiency. Experiments are conducted on various challenging cross-dataset distribution shift scenarios, including nuScenes, Lyft, and Waymo. Results demonstrate that our method achieves superior adaptation accuracy, surpassing state-of-the-art test-time training methods for trajectory prediction. Additionally, our method excels under suboptimal learning rates and high FPS demands, showcasing its robustness and practicality.

cs.CV

Impact of Dataset Properties on Membership Inference Vulnerability of Deep Transfer Learning

Membership inference attacks (MIAs) are used to test practical privacy of machine learning models. MIAs complement formal guarantees from differential privacy (DP) under a more realistic adversary model. We analyse MIA vulnerability of fine-tuned neural networks both empirically and theoretically, the latter using a simplified model of fine-tuning. We show that the vulnerability of non-DP models when measured as the attacker advantage at a fixed false positive rate reduces according to a simple power law as the number of examples per class increases. A similar power-law applies even for the most vulnerable points, but the dataset size needed for adequate protection of the most vulnerable points is very large.

cs.CR

What Breaks Knowledge Graph based RAG? Benchmarking and Empirical Insights into Reasoning under Incomplete Knowledge

Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) is an increasingly explored approach for combining the reasoning capabilities of large language models with the structured evidence of knowledge graphs. However, current evaluation practices fall short: existing benchmarks often include questions that can be directly answered using existing triples in KG, making it unclear whether models perform reasoning or simply retrieve answers directly. Moreover, inconsistent evaluation metrics and lenient answer matching criteria further obscure meaningful comparisons. In this work, we introduce a general method for constructing benchmarks and present BRINK (Benchmark for Reasoning under Incomplete Knowledge) to systematically assess KG-RAG methods under knowledge incompleteness. Our empirical results show that current KG-RAG methods have limited reasoning ability under missing knowledge, often rely on internal memorization, and exhibit varying degrees of generalization depending on their design.

cs.AI

Necking of epithelial tissues with cellular topological transition

As the cover of embryos and adult organisms, epithelial tissues are subjected to substantial mechanical forces in tissue morphogenesis. However, the finite deformation behaviors of epithelial tissues remain largely unexplored. This study combines discrete vertex simulations with a multiscale constitutive model to investigate the necking behavior of epithelial tissues. In the multiscale model, the shape changes and topological transitions of single cells are mapped to the elastic and inelastic tissue deformations via a mean-field formulation. Our results show that the necking bifurcation of a stretched tissue arises from cellular topological transitions. The bifurcation condition and the steady state of necking propagation are predicted from the constitutive model and validated by vertex simulations. Furthermore, we find that topological defects in disordered tissues facilitate necking bifurcation but impede its propagation. These defects also induce the necked region to collapse into a thin thread, as observed in real tissues. Together, our work provides valuable insights into the deformation behaviors of epithelial tissues.

physics.bio-ph

GR-Agent: Adaptive Graph Reasoning Agent under Incomplete Knowledge

Large language models (LLMs) achieve strong results on knowledge graph question answering (KGQA), but most benchmarks assume complete knowledge graphs (KGs) where direct supporting triples exist. This reduces evaluation to shallow retrieval and overlooks the reality of incomplete KGs, where many facts are missing and answers must be inferred from existing facts. We bridge this gap by proposing a methodology for constructing benchmarks under KG incompleteness, which removes direct supporting triples while ensuring that alternative reasoning paths required to infer the answer remain. Experiments on benchmarks constructed using our methodology show that existing methods suffer consistent performance degradation under incompleteness, highlighting their limited reasoning ability. To overcome this limitation, we present the Adaptive Graph Reasoning Agent (GR-Agent). It first constructs an interactive environment from the KG, and then formalizes KGQA as agent environment interaction within this environment. GR-Agent operates over an action space comprising graph reasoning tools and maintains a memory of potential supporting reasoning evidence, including relevant relations and reasoning paths. Extensive experiments demonstrate that GR-Agent outperforms non-training baselines and performs comparably to training-based methods under both complete and incomplete settings.

cs.AI

Conceptual Design of the Muonium-to-Antimuonium Conversion Experiment (MACE)

The spontaneous conversion of muonium to antimuonium is one of the interesting charged lepton flavor violation phenomena offering a sensitive probe of potential new physics and serving as a tool to constrain the parameter space beyond the Standard Model. The Muonium-to-Antimuonium Conversion Experiment (MACE) is designed to utilize a high-intensity muon beam, a Michel electron magnetic spectrometer, a positron transport system, and a positron detection system, to either discover or constrain this rare process with a conversion probability of $\mathcal{O}(10^{-13})$. This article presents an overview of the theoretical framework as well as a detailed description of the experimental design for the search for muonium-to-antimuonium conversion.

hep-ex

Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning

Forecasting future links is a central task in temporal graph (TG) reasoning, requiring models to leverage historical interactions to predict upcoming ones. Traditional neural approaches, such as temporal graph neural networks, achieve strong performance but lack explainability and cannot be applied to unseen graphs without retraining. Recent studies have begun to explore using large language models (LLMs) for graph reasoning, but most of them are constrained to static graphs or small synthetic TGs and lack the evaluation of the quality of reasoning traces generated by LLMs. In this work, we present Reasoning-Enhanced Learning for Temporal Graphs (ReaL-TG), a reinforcement learning framework that fine-tunes LLMs to perform explainable link forecasting on real-world TGs. ReaL-TG uses outcome-based reward to encourage models to self-explore reasoning strategies from graph structure and to produce explanations that directly justify their predictions. To enable evaluation on LLM-generated reasoning traces, we propose a new evaluation protocol combining ranking metrics with an LLM-as-a-Judge system that assesses both the quality of reasoning and the impact of hallucinations. Experiments with ReaL-TG-4B, obtained by fine-tuning Qwen3-4B under our framework, show that it outperforms much larger frontier LLMs, including GPT-5 mini, on ranking metrics, while producing high-quality explanations confirmed by both the LLM judge and human evaluation.

cs.AI